查找两个文本文件的差异
Find differences of two text files
我有两个超过 100000 行的单独文件 A
和 B
。现在我需要比较每一行以查找它们是否在 A
中但不在 B
中。两个文件都是文本格式。
文件A:
>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP
文件 B:
>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP
我需要:A-B
>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
如有任何帮助或建议,我们将不胜感激。
你可以试试regular expression
import re
# here I am taking as text but u can read the file like text_1 = open('file_1.txt').read()
text_1 = """>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP"""
text_2 = """>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP
"""
protein_1 = {i[0]:i[1] for i in re.findall(r'^>Q63544|(\d+)\n(.*)', text_1, re.MULTILINE) if len(i[0])>0}
protein_2 = {i[0]:i[1] for i in re.findall(r'^>Q63544|(\d+)\n(.*)', text_2, re.MULTILINE) if len(i[0])>0}
dissimilar_protein = {i:protein_1[i] for i in protein_1 if i not in protein_2}
print(list(dissimilar_protein.values()))
['----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY', 'TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP']
或者如果您想要特定格式
output = [f'>Q63544|{i}\n{j}' for i,j in dissimilar_protein.items()]
for i in output:
print(i)
>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
NOTR:我假设所有蛋白质编号都以 Q63544
开头
我有两个超过 100000 行的单独文件 A
和 B
。现在我需要比较每一行以查找它们是否在 A
中但不在 B
中。两个文件都是文本格式。
文件A:
>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP
文件 B:
>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP
我需要:A-B
>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
如有任何帮助或建议,我们将不胜感激。
你可以试试regular expression
import re
# here I am taking as text but u can read the file like text_1 = open('file_1.txt').read()
text_1 = """>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP"""
text_2 = """>Q63544|51
KTKQGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKT
>Q63544|54
QGVTEAAEKTKEGVMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEE
>Q63544|67
VMYVGTKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRK
>Q63544|73
KTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEPP
"""
protein_1 = {i[0]:i[1] for i in re.findall(r'^>Q63544|(\d+)\n(.*)', text_1, re.MULTILINE) if len(i[0])>0}
protein_2 = {i[0]:i[1] for i in re.findall(r'^>Q63544|(\d+)\n(.*)', text_2, re.MULTILINE) if len(i[0])>0}
dissimilar_protein = {i:protein_1[i] for i in protein_1 if i not in protein_2}
print(list(dissimilar_protein.values()))
['----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY', 'TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP']
或者如果您想要特定格式
output = [f'>Q63544|{i}\n{j}' for i,j in dissimilar_protein.items()]
for i in output:
print(i)
>Q63544|9
----------------------MDVFKKGFSIAREGVVGAVEKTKQGVTEAAEKTKEGVMY
>Q63544|72
TKTKGERGTSVTSVAEKTKEQANAVSEAVVSSVNTVATKTVEEAENIVVTTGVVRKEDLEP
NOTR:我假设所有蛋白质编号都以 Q63544