Sequences¶
|
Single biological sequence with a header |
|
Alphabet of arbitrary biological sequences |
Immutable container with commonly used biological sequence alphabets |
|
|
Guess the alphabet of a sequence string: the alphabet that contains the most sequence characters. |
|
Collection of (unaligned) DNA or amino acid sequences |
|
Collection of aligned DNA or amino acid sequences, stored as a numpy matrix with one row per sequence. |
|
(Partially) abstract base class for sequence collections. |
|
Iterator over the sequences in a fasta or json formatted file or string. |
|
Iterator over batches of sequences in a fasta or json formatted file or string. |
- class picea.Sequence(header=None, sequence=None, alphabet=None, annotation=<factory>)[source]¶
Single biological sequence with a header
Examples
>>> dna = Sequence('test_dna', 'ACGATCGACTAGCA') >>> dna.alphabet.name 'DNA' >>> protein = Sequence('test_aa', 'QAPISAIWPOIWQ*') >>> protein.alphabet.name 'AminoAcid' >>> dna.reverse_complement.sequence 'TGCTAGTCGATCGT'
- Parameters:
header (str) – Sequence name
sequence (str) – Sequence string
alphabet (Alphabet) – Sequence alphabet. Detected from the sequence when not given (see
guess_alphabet()). Sequences derived from this one (slices, complements, etc.) keep the alphabet.annotation (Optional[SequenceAnnotation]) – Annotation of the sequence. Defaults to an empty annotation.
- property uppercase: Sequence¶
All sequence characters in uppercase
Examples
>>> s = Sequence('test_dna', 'acgTA') >>> s.uppercase.sequence 'ACGTA'
- property lowercase: Sequence¶
All sequence characters in lowercase
Examples
>>> s = Sequence('test_dna', 'acgTA') >>> s.lowercase.sequence 'acgta'
- property reverse: Sequence¶
Reverse sequence order
Examples
>>> s = Sequence('test_dna', 'ACGTA') >>> s.reverse.sequence 'ATGCA'
- property complement: Sequence¶
Complement DNA sequences based on watson-crick pairing
Examples
>>> s = Sequence('test_dna', 'ACGTA') >>> s.complement.sequence 'TGCAT'
- property reverse_complement: Sequence¶
Reverse sequence order and complement nucleotides vased on watson-crick pairing. This is the same as accessing the reversed and then complemented sequence (in arbitrary order)
Examples
>>> s = Sequence('test_dna', 'ACGTA') >>> s.reverse_complement.sequence 'TACGT' >>> s.reverse_complement.sequence == s.reverse.complement.sequence == s.complement.reverse.sequence True
- property amino_acids: Sequence¶
Translate nucleotide codon triplets into amino acids
Examples
>>> s = Sequence('test_dna', 'ATGATGTAA') >>> s.amino_acids.sequence 'MM*'
- to_dict()[source]¶
Make dictionary with header and sequence elements
Examples
>>> s = Sequence('test', 'ACGTA') >>> s.to_dict() {'header': 'test', 'sequence': 'ACGTA'}
- class picea.Alphabet(name, members)[source]¶
Alphabet of arbitrary biological sequences
Examples
>>> DNA = Alphabet('DNA', 'ACGT') >>> DNA Alphabet(name='DNA', members='ACGT')
>>> Protein = Alphabet('AminoAcid', '*-?ACDEFGHIKLMNPQRSTVWXY') >>> Protein Alphabet(name='AminoAcid', members='*-?ACDEFGHIKLMNPQRSTVWXY')
- score(sequence, match=1.0, mismatch=-1.0, n_chars=100)[source]¶
Scores how well a sequence matches an alphabet by summing (mis)matches of sequence letters that are not in the alphabet and (mis)matches of alphabet letters that are not in the sequence.
- Parameters:
sequence (str) – Sequence string for which to determine how well it fits the alphabet
match (float, optional) – match score. Defaults to 1.0.
mismatch (float, optional) – mismatch score. Defaults to -1.0.
n_chars (int, optional) – number of sequence characters to use in scoring. Large numbers incur a significant computational cost.
- Returns:
Score of how well a sequence matches the alphabet
- Return type:
(float)
- picea.alphabets¶
Immutable container with commonly used biological sequence alphabets
- picea.guess_alphabet(sequence, n_chars=100)[source]¶
Guess the alphabet of a sequence string: the alphabet that contains the most sequence characters. Ties go to the smallest alphabet, so a sequence of only nucleotide characters is DNA.
Examples
>>> guess_alphabet('ACGTTGCA').name 'DNA' >>> guess_alphabet('MKVLAAGIVGLL').name 'AminoAcid'
- class picea.SequenceCollection(sequences=None, sequence_annotation=None)[source]¶
Collection of (unaligned) DNA or amino acid sequences
Examples
>>> seqs = SequenceCollection([('a', 'ACGT'), ('b', 'GGCCAA')]) >>> len(seqs) 2 >>> seqs['b'].sequence 'GGCCAA'
- Parameters:
sequences (Iterable[Tuple[str, str]]) – (header, sequence) tuples
sequence_annotation (SequenceAnnotation) – Annotation of the sequences
- property headers: List[str]¶
Sequence headers, in insertion order
- Returns:
Sequence headers
- Return type:
List[str]
- property n_seqs: int¶
Number of sequences in the collection
- Returns:
Number of sequences
- Return type:
- align(method='mafft', method_kwargs=None)[source]¶
Align the sequences with an external multiple sequence aligner. The aligner is called as
<method> <method_kwargs> -, and must read fasta from stdin and write aligned fasta to stdout (like MAFFT).- Parameters:
- Returns:
Aligned sequences
- Return type:
- Raises:
RuntimeError – If the aligner exits with an error
- classmethod from_fasta(filename=None, string=None)¶
Read a fasta formatted file or string. Exactly one of
filenameorstringmust be given.Examples
>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC') >>> seqs.headers ['a', 'b']
- Parameters:
- Returns:
New collection of the class this method is called on
- Return type:
- classmethod from_json(filename=None, string=None)¶
Read a json formatted file or string, as written by
to_json(). Exactly one offilenameorstringmust be given.- Parameters:
- Returns:
New collection of the class this method is called on
- Return type:
- classmethod from_sequence_iter(sequence_iter)¶
Create a collection from
Sequenceobjects- Parameters:
sequence_iter (Iterable[Sequence]) – Sequences
- Returns:
New collection of the class this method is called on
- Return type:
- property iloc: SequenceIndex¶
index with an int, a list of ints, or a slice to get a new collection of the same type with the selected sequences
Examples
>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC\n>c\nTTAA') >>> seqs.iloc[1:].headers ['b', 'c'] >>> seqs.iloc[[0, 2]].headers ['a', 'c']
- Type:
Position based indexing
- modify_inplace(mod_func)¶
Change headers and/or sequences in place by calling
mod_funcon every sequenceExamples
>>> seqs = SequenceCollection.from_fasta(string='>a\nacgt\n>b\nggcc') >>> seqs.modify_inplace(lambda header, sequence: (header.upper(), sequence.upper())) >>> seqs.to_fasta() '>A\nACGT\n>B\nGGCC'
- rename_inplace(rename_func)¶
Rename all headers in place by calling
rename_funcon every header
- property sequences: List[str]¶
List of sequences without headers
- Returns:
list of sequences
- Return type:
List[str]
- to_fasta(linewidth=80)¶
fasta formatted string of all sequences
- to_json(indent=None)¶
json formatted string: a list of objects with
headerandsequencekeys- Parameters:
indent (Optional[int]) – Indentation, passed to
json.dumps()- Returns:
json formatted string
- Return type:
- class picea.MultipleSequenceAlignment(sequences=None, sequence_annotation=None)[source]¶
Collection of aligned DNA or amino acid sequences, stored as a numpy matrix with one row per sequence. Sequences shorter than the alignment are padded with gaps (
-) at the end.Examples
>>> msa = MultipleSequenceAlignment.from_fasta(string='>a\nAC-GT\n>b\nACCGT') >>> msa.shape (2, 5)
- Parameters:
sequences (Optional[Iterable[Sequence]]) – Sequences
sequence_annotation (Optional[SequenceAnnotation]) – Annotation of the sequences
- property headers: List[str]¶
Sequence headers, in insertion order
- Returns:
Sequence headers
- Return type:
List[str]
- property n_seqs: int¶
Number of sequences in the collection
- Returns:
Number of sequences
- Return type:
- to_nexus()[source]¶
Nexus
datablock with the alignment. The datatype isproteinif any sequence is an amino acid sequence, anddnaotherwise.Examples
>>> msa = MultipleSequenceAlignment.from_fasta(string='>a\nAC-GT\n>b\nACCGT') >>> print(msa.to_nexus()) begin data; dimensions ntax=2 nchar=5; format datatype=dna gap=-; matrix a AC-GT b ACCGT ; end;
- Returns:
Nexus formatted string
- Return type:
- align(method='mafft', method_kwargs=None)¶
Align the sequences with an external multiple sequence aligner. The aligner is called as
<method> <method_kwargs> -, and must read fasta from stdin and write aligned fasta to stdout (like MAFFT).- Parameters:
- Returns:
Aligned sequences
- Return type:
- Raises:
RuntimeError – If the aligner exits with an error
- classmethod from_fasta(filename=None, string=None)¶
Read a fasta formatted file or string. Exactly one of
filenameorstringmust be given.Examples
>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC') >>> seqs.headers ['a', 'b']
- Parameters:
- Returns:
New collection of the class this method is called on
- Return type:
- classmethod from_json(filename=None, string=None)¶
Read a json formatted file or string, as written by
to_json(). Exactly one offilenameorstringmust be given.- Parameters:
- Returns:
New collection of the class this method is called on
- Return type:
- classmethod from_sequence_iter(sequence_iter)¶
Create a collection from
Sequenceobjects- Parameters:
sequence_iter (Iterable[Sequence]) – Sequences
- Returns:
New collection of the class this method is called on
- Return type:
- property iloc: SequenceIndex¶
index with an int, a list of ints, or a slice to get a new collection of the same type with the selected sequences
Examples
>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC\n>c\nTTAA') >>> seqs.iloc[1:].headers ['b', 'c'] >>> seqs.iloc[[0, 2]].headers ['a', 'c']
- Type:
Position based indexing
- modify_inplace(mod_func)¶
Change headers and/or sequences in place by calling
mod_funcon every sequenceExamples
>>> seqs = SequenceCollection.from_fasta(string='>a\nacgt\n>b\nggcc') >>> seqs.modify_inplace(lambda header, sequence: (header.upper(), sequence.upper())) >>> seqs.to_fasta() '>A\nACGT\n>B\nGGCC'
- rename_inplace(rename_func)¶
Rename all headers in place by calling
rename_funcon every header
- property sequences: List[str]¶
List of sequences without headers
- Returns:
list of sequences
- Return type:
List[str]
- to_fasta(linewidth=80)¶
fasta formatted string of all sequences
- to_json(indent=None)¶
json formatted string: a list of objects with
headerandsequencekeys- Parameters:
indent (Optional[int]) – Indentation, passed to
json.dumps()- Returns:
json formatted string
- Return type:
- class picea.AbstractSequenceCollection(sequences=None, sequence_annotation=None)[source]¶
(Partially) abstract base class for sequence collections.
Subclasses implement storage:
__setitem__,__getitem__,__delitem__,pop(),headersandn_seqs. All other methods (reading and writing fasta and json, iteration, indexing withiloc, renaming) build on these.Sequences are stored by header, and indexing with a header returns a
Sequence. Setting a header that already exists does not overwrite the existing sequence: the new sequence gets a unique header (header_1,header_2, …) and a warning is issued.- abstract property headers: List[str]¶
Sequence headers, in insertion order
- Returns:
Sequence headers
- Return type:
List[str]
- property iloc: SequenceIndex¶
index with an int, a list of ints, or a slice to get a new collection of the same type with the selected sequences
Examples
>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC\n>c\nTTAA') >>> seqs.iloc[1:].headers ['b', 'c'] >>> seqs.iloc[[0, 2]].headers ['a', 'c']
- Type:
Position based indexing
- property sequences: List[str]¶
List of sequences without headers
- Returns:
list of sequences
- Return type:
List[str]
- abstract property n_seqs: int¶
Number of sequences in the collection
- Returns:
Number of sequences
- Return type:
- classmethod from_sequence_iter(sequence_iter)[source]¶
Create a collection from
Sequenceobjects- Parameters:
sequence_iter (Iterable[Sequence]) – Sequences
- Returns:
New collection of the class this method is called on
- Return type:
- classmethod from_fasta(filename=None, string=None)[source]¶
Read a fasta formatted file or string. Exactly one of
filenameorstringmust be given.Examples
>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC') >>> seqs.headers ['a', 'b']
- Parameters:
- Returns:
New collection of the class this method is called on
- Return type:
- classmethod from_json(filename=None, string=None)[source]¶
Read a json formatted file or string, as written by
to_json(). Exactly one offilenameorstringmust be given.- Parameters:
- Returns:
New collection of the class this method is called on
- Return type:
- to_json(indent=None)[source]¶
json formatted string: a list of objects with
headerandsequencekeys- Parameters:
indent (Optional[int]) – Indentation, passed to
json.dumps()- Returns:
json formatted string
- Return type:
- add(seq)[source]¶
Add a sequence to the collection (in place)
- Parameters:
seq (Sequence) – Sequence to add
- modify_inplace(mod_func)[source]¶
Change headers and/or sequences in place by calling
mod_funcon every sequenceExamples
>>> seqs = SequenceCollection.from_fasta(string='>a\nacgt\n>b\nggcc') >>> seqs.modify_inplace(lambda header, sequence: (header.upper(), sequence.upper())) >>> seqs.to_fasta() '>A\nACGT\n>B\nGGCC'
- class picea.SequenceReader(string=None, filename=None, filetype=None)[source]¶
Iterator over the sequences in a fasta or json formatted file or string. Exactly one of
stringorfilenamemust be given. json input is a list of objects withheaderandsequencekeys.Examples
>>> fasta_string = '>1\nACGC\n>2\nTGTGTA\n' >>> [seq.header for seq in SequenceReader(string=fasta_string, filetype='fasta')] ['1', '2']
- Parameters:
- Raises:
ValueError – If
filetypeis not supported
- class picea.BatchSequenceReader(string=None, filename=None, filetype=None, batchsize=10)[source]¶
Iterator over batches of sequences in a fasta or json formatted file or string. Every batch is a
SequenceCollectionofbatchsizesequences, except for the last batch, which can be smaller.Examples
>>> fasta_string = '>1\nACGC\n>2\nTGTGTA\n>3\nAAGT\n>4\nCCA\n>5\nGGA\n' >>> reader = BatchSequenceReader(string=fasta_string, filetype='fasta', batchsize=2) >>> [batch.headers for batch in reader] [['1', '2'], ['3', '4'], ['5']]