Sequences

Sequence([header, sequence, alphabet, ...])

Single biological sequence with a header

Alphabet(name, members)

Alphabet of arbitrary biological sequences

alphabets

Immutable container with commonly used biological sequence alphabets

guess_alphabet(sequence[, n_chars])

Guess the alphabet of a sequence string: the alphabet that contains the most sequence characters.

SequenceCollection([sequences, ...])

Collection of (unaligned) DNA or amino acid sequences

MultipleSequenceAlignment([sequences, ...])

Collection of aligned DNA or amino acid sequences, stored as a numpy matrix with one row per sequence.

AbstractSequenceCollection([sequences, ...])

(Partially) abstract base class for sequence collections.

SequenceReader([string, filename, filetype])

Iterator over the sequences in a fasta or json formatted file or string.

BatchSequenceReader([string, filename, ...])

Iterator over batches of sequences in a fasta or json formatted file or string.

class picea.Sequence(header=None, sequence=None, alphabet=None, annotation=<factory>)[source]

Single biological sequence with a header

Examples

>>> dna = Sequence('test_dna', 'ACGATCGACTAGCA')
>>> dna.alphabet.name
'DNA'
>>> protein = Sequence('test_aa', 'QAPISAIWPOIWQ*')
>>> protein.alphabet.name
'AminoAcid'
>>> dna.reverse_complement.sequence
'TGCTAGTCGATCGT'
Parameters:
  • header (str) – Sequence name

  • sequence (str) – Sequence string

  • alphabet (Alphabet) – Sequence alphabet. Detected from the sequence when not given (see guess_alphabet()). Sequences derived from this one (slices, complements, etc.) keep the alphabet.

  • annotation (Optional[SequenceAnnotation]) – Annotation of the sequence. Defaults to an empty annotation.

property uppercase: Sequence

All sequence characters in uppercase

Examples

>>> s = Sequence('test_dna', 'acgTA')
>>> s.uppercase.sequence
'ACGTA'
property lowercase: Sequence

All sequence characters in lowercase

Examples

>>> s = Sequence('test_dna', 'acgTA')
>>> s.lowercase.sequence
'acgta'
property reverse: Sequence

Reverse sequence order

Examples

>>> s = Sequence('test_dna', 'ACGTA')
>>> s.reverse.sequence
'ATGCA'
property complement: Sequence

Complement DNA sequences based on watson-crick pairing

Examples

>>> s = Sequence('test_dna', 'ACGTA')
>>> s.complement.sequence
'TGCAT'
property reverse_complement: Sequence

Reverse sequence order and complement nucleotides vased on watson-crick pairing. This is the same as accessing the reversed and then complemented sequence (in arbitrary order)

Examples

>>> s = Sequence('test_dna', 'ACGTA')
>>> s.reverse_complement.sequence
'TACGT'
>>> s.reverse_complement.sequence == s.reverse.complement.sequence == s.complement.reverse.sequence
True
property amino_acids: Sequence

Translate nucleotide codon triplets into amino acids

Examples

>>> s = Sequence('test_dna', 'ATGATGTAA')
>>> s.amino_acids.sequence
'MM*'
to_dict()[source]

Make dictionary with header and sequence elements

Examples

>>> s = Sequence('test', 'ACGTA')
>>> s.to_dict()
{'header': 'test', 'sequence': 'ACGTA'}
Returns:

sequence dictionary

Return type:

Dict[str, str]

classmethod from_fasta(string)[source]

Create a sequence from a fasta formatted string (single sequence only)

Examples

>>> s = Sequence.from_fasta('>test\nACGT')
>>> s.header, s.sequence
('test', 'ACGT')
Parameters:

string (str) – fasta formatted string

Returns:

Sequence

Return type:

Sequence

to_fasta(linewidth=80)[source]

fasta formatted string

Examples

>>> s = Sequence('test_dna', 'ACGTA')
>>> s.to_fasta()
'>test_dna\nACGTA'
Parameters:

linewidth (int) – Maximum number of sequence characters per line

Returns:

fasta formatted string

Return type:

str

class picea.Alphabet(name, members)[source]

Alphabet of arbitrary biological sequences

Examples

>>> DNA = Alphabet('DNA', 'ACGT')
>>> DNA
Alphabet(name='DNA', members='ACGT')
>>> Protein = Alphabet('AminoAcid', '*-?ACDEFGHIKLMNPQRSTVWXY')
>>> Protein
Alphabet(name='AminoAcid', members='*-?ACDEFGHIKLMNPQRSTVWXY')
Parameters:
  • name (str) – Alphabet name

  • members (Iterable[str]) – Letters of the alphabet

score(sequence, match=1.0, mismatch=-1.0, n_chars=100)[source]

Scores how well a sequence matches an alphabet by summing (mis)matches of sequence letters that are not in the alphabet and (mis)matches of alphabet letters that are not in the sequence.

Parameters:
  • sequence (str) – Sequence string for which to determine how well it fits the alphabet

  • match (float, optional) – match score. Defaults to 1.0.

  • mismatch (float, optional) – mismatch score. Defaults to -1.0.

  • n_chars (int, optional) – number of sequence characters to use in scoring. Large numbers incur a significant computational cost.

Returns:

Score of how well a sequence matches the alphabet

Return type:

(float)

validate(sequence)[source]

Determine whether a sequence strictly fits an alphabet

Parameters:

sequence (str) – Sequence string

Returns:

true if all characters in sequence are in the alphabet

Return type:

bool

complement(sequence)[source]

Returns complementary strand of DNA or RNA sequence strings

Examples

>>> DNA = Alphabet('DNA', 'ACGT')
>>> DNA.complement('AACTACG')
'TTGATGC'
Parameters:

sequence (str) – Sequence string

Returns:

complementary strand sequence string

Return type:

str

translate(sequence)[source]

Translate DNA or RNA sequence string to amino acid string

Examples

>>> DNA = Alphabet('DNA', 'ACGT')
>>> DNA.translate('ATGACGACGTAA')
'MTT*'
Parameters:

sequence (str) – Sequence string (sequence length must be multiple of 3)

Returns:

Amino acid string

Return type:

str

picea.alphabets

Immutable container with commonly used biological sequence alphabets

picea.guess_alphabet(sequence, n_chars=100)[source]

Guess the alphabet of a sequence string: the alphabet that contains the most sequence characters. Ties go to the smallest alphabet, so a sequence of only nucleotide characters is DNA.

Examples

>>> guess_alphabet('ACGTTGCA').name
'DNA'
>>> guess_alphabet('MKVLAAGIVGLL').name
'AminoAcid'
Parameters:
  • sequence (str) – Sequence string

  • n_chars (int, optional) – Number of sequence characters to use. Defaults to 100.

Returns:

Best matching alphabet

Return type:

Alphabet

class picea.SequenceCollection(sequences=None, sequence_annotation=None)[source]

Collection of (unaligned) DNA or amino acid sequences

Examples

>>> seqs = SequenceCollection([('a', 'ACGT'), ('b', 'GGCCAA')])
>>> len(seqs)
2
>>> seqs['b'].sequence
'GGCCAA'
Parameters:
  • sequences (Iterable[Tuple[str, str]]) – (header, sequence) tuples

  • sequence_annotation (SequenceAnnotation) – Annotation of the sequences

property headers: List[str]

Sequence headers, in insertion order

Returns:

Sequence headers

Return type:

List[str]

property n_seqs: int

Number of sequences in the collection

Returns:

Number of sequences

Return type:

int

align(method='mafft', method_kwargs=None)[source]

Align the sequences with an external multiple sequence aligner. The aligner is called as <method> <method_kwargs> -, and must read fasta from stdin and write aligned fasta to stdout (like MAFFT).

Parameters:
  • method (Optional[str]) – Name of the aligner executable

  • method_kwargs (Optional[Mapping[str, str]]) – Command line options for the aligner, e.g. {"--thread": "4"}

Returns:

Aligned sequences

Return type:

MultipleSequenceAlignment

Raises:

RuntimeError – If the aligner exits with an error

pop(header)[source]

Remove a sequence from the collection and return it

Parameters:

header (str) – Header of the sequence to remove

Returns:

The removed sequence

Return type:

Sequence

Raises:

KeyError – If there is no sequence with this header

add(seq)

Add a sequence to the collection (in place)

Parameters:

seq (Sequence) – Sequence to add

classmethod from_fasta(filename=None, string=None)

Read a fasta formatted file or string. Exactly one of filename or string must be given.

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC')
>>> seqs.headers
['a', 'b']
Parameters:
  • filename (str) – fasta filename

  • string (str) – fasta formatted string

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

classmethod from_json(filename=None, string=None)

Read a json formatted file or string, as written by to_json(). Exactly one of filename or string must be given.

Parameters:
  • filename (Optional[str]) – json filename

  • string (Optional[str]) – json formatted string: a list of objects with header and sequence keys

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

classmethod from_sequence_iter(sequence_iter)

Create a collection from Sequence objects

Parameters:

sequence_iter (Iterable[Sequence]) – Sequences

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

property iloc: SequenceIndex

index with an int, a list of ints, or a slice to get a new collection of the same type with the selected sequences

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC\n>c\nTTAA')
>>> seqs.iloc[1:].headers
['b', 'c']
>>> seqs.iloc[[0, 2]].headers
['a', 'c']
Type:

Position based indexing

modify_inplace(mod_func)

Change headers and/or sequences in place by calling mod_func on every sequence

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nacgt\n>b\nggcc')
>>> seqs.modify_inplace(lambda header, sequence: (header.upper(), sequence.upper()))
>>> seqs.to_fasta()
'>A\nACGT\n>B\nGGCC'
Parameters:

mod_func (Callable[[str, str], tuple[str, str]]) – Function that takes a header and a sequence string, and returns a new header and sequence string

rename_inplace(rename_func)

Rename all headers in place by calling rename_func on every header

Parameters:

rename_func (Callable[[str], str]) – Function that takes a header and returns a new header

property sequences: List[str]

List of sequences without headers

Returns:

list of sequences

Return type:

List[str]

to_fasta(linewidth=80)

fasta formatted string of all sequences

Parameters:

linewidth (int) – Maximum number of sequence characters per line

Returns:

Multi-line fasta formatted string

Return type:

str

to_json(indent=None)

json formatted string: a list of objects with header and sequence keys

Parameters:

indent (Optional[int]) – Indentation, passed to json.dumps()

Returns:

json formatted string

Return type:

str

class picea.MultipleSequenceAlignment(sequences=None, sequence_annotation=None)[source]

Collection of aligned DNA or amino acid sequences, stored as a numpy matrix with one row per sequence. Sequences shorter than the alignment are padded with gaps (-) at the end.

Examples

>>> msa = MultipleSequenceAlignment.from_fasta(string='>a\nAC-GT\n>b\nACCGT')
>>> msa.shape
(2, 5)
Parameters:
  • sequences (Optional[Iterable[Sequence]]) – Sequences

  • sequence_annotation (Optional[SequenceAnnotation]) – Annotation of the sequences

property headers: List[str]

Sequence headers, in insertion order

Returns:

Sequence headers

Return type:

List[str]

property n_seqs: int

Number of sequences in the collection

Returns:

Number of sequences

Return type:

int

property n_chars: int

Number of alignment columns

property shape: Tuple[int, int]

Number of sequences and number of alignment columns

to_nexus()[source]

Nexus data block with the alignment. The datatype is protein if any sequence is an amino acid sequence, and dna otherwise.

Examples

>>> msa = MultipleSequenceAlignment.from_fasta(string='>a\nAC-GT\n>b\nACCGT')
>>> print(msa.to_nexus())
begin data;
    dimensions ntax=2 nchar=5;
    format datatype=dna gap=-;
    matrix
    a AC-GT
    b ACCGT
    ;
end;
Returns:

Nexus formatted string

Return type:

str

add(seq)

Add a sequence to the collection (in place)

Parameters:

seq (Sequence) – Sequence to add

align(method='mafft', method_kwargs=None)

Align the sequences with an external multiple sequence aligner. The aligner is called as <method> <method_kwargs> -, and must read fasta from stdin and write aligned fasta to stdout (like MAFFT).

Parameters:
  • method (Optional[str]) – Name of the aligner executable

  • method_kwargs (Optional[Mapping[str, str]]) – Command line options for the aligner, e.g. {"--thread": "4"}

Returns:

Aligned sequences

Return type:

MultipleSequenceAlignment

Raises:

RuntimeError – If the aligner exits with an error

classmethod from_fasta(filename=None, string=None)

Read a fasta formatted file or string. Exactly one of filename or string must be given.

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC')
>>> seqs.headers
['a', 'b']
Parameters:
  • filename (str) – fasta filename

  • string (str) – fasta formatted string

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

classmethod from_json(filename=None, string=None)

Read a json formatted file or string, as written by to_json(). Exactly one of filename or string must be given.

Parameters:
  • filename (Optional[str]) – json filename

  • string (Optional[str]) – json formatted string: a list of objects with header and sequence keys

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

classmethod from_sequence_iter(sequence_iter)

Create a collection from Sequence objects

Parameters:

sequence_iter (Iterable[Sequence]) – Sequences

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

property iloc: SequenceIndex

index with an int, a list of ints, or a slice to get a new collection of the same type with the selected sequences

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC\n>c\nTTAA')
>>> seqs.iloc[1:].headers
['b', 'c']
>>> seqs.iloc[[0, 2]].headers
['a', 'c']
Type:

Position based indexing

modify_inplace(mod_func)

Change headers and/or sequences in place by calling mod_func on every sequence

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nacgt\n>b\nggcc')
>>> seqs.modify_inplace(lambda header, sequence: (header.upper(), sequence.upper()))
>>> seqs.to_fasta()
'>A\nACGT\n>B\nGGCC'
Parameters:

mod_func (Callable[[str, str], tuple[str, str]]) – Function that takes a header and a sequence string, and returns a new header and sequence string

pop(header)[source]

Remove a sequence from the collection and return it

Parameters:

header (str) – Header of the sequence to remove

Returns:

The removed sequence

Return type:

Sequence

Raises:

KeyError – If there is no sequence with this header

rename_inplace(rename_func)

Rename all headers in place by calling rename_func on every header

Parameters:

rename_func (Callable[[str], str]) – Function that takes a header and returns a new header

property sequences: List[str]

List of sequences without headers

Returns:

list of sequences

Return type:

List[str]

to_fasta(linewidth=80)

fasta formatted string of all sequences

Parameters:

linewidth (int) – Maximum number of sequence characters per line

Returns:

Multi-line fasta formatted string

Return type:

str

to_json(indent=None)

json formatted string: a list of objects with header and sequence keys

Parameters:

indent (Optional[int]) – Indentation, passed to json.dumps()

Returns:

json formatted string

Return type:

str

pairwise_distances(distance_measure='identity')[source]

Not implemented yet, returns None

class picea.AbstractSequenceCollection(sequences=None, sequence_annotation=None)[source]

(Partially) abstract base class for sequence collections.

Subclasses implement storage: __setitem__, __getitem__, __delitem__, pop(), headers and n_seqs. All other methods (reading and writing fasta and json, iteration, indexing with iloc, renaming) build on these.

Sequences are stored by header, and indexing with a header returns a Sequence. Setting a header that already exists does not overwrite the existing sequence: the new sequence gets a unique header (header_1, header_2, …) and a warning is issued.

abstract property headers: List[str]

Sequence headers, in insertion order

Returns:

Sequence headers

Return type:

List[str]

property iloc: SequenceIndex

index with an int, a list of ints, or a slice to get a new collection of the same type with the selected sequences

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC\n>c\nTTAA')
>>> seqs.iloc[1:].headers
['b', 'c']
>>> seqs.iloc[[0, 2]].headers
['a', 'c']
Type:

Position based indexing

property sequences: List[str]

List of sequences without headers

Returns:

list of sequences

Return type:

List[str]

abstract property n_seqs: int

Number of sequences in the collection

Returns:

Number of sequences

Return type:

int

classmethod from_sequence_iter(sequence_iter)[source]

Create a collection from Sequence objects

Parameters:

sequence_iter (Iterable[Sequence]) – Sequences

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

classmethod from_fasta(filename=None, string=None)[source]

Read a fasta formatted file or string. Exactly one of filename or string must be given.

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nACGT\n>b\nGGCC')
>>> seqs.headers
['a', 'b']
Parameters:
  • filename (str) – fasta filename

  • string (str) – fasta formatted string

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

to_fasta(linewidth=80)[source]

fasta formatted string of all sequences

Parameters:

linewidth (int) – Maximum number of sequence characters per line

Returns:

Multi-line fasta formatted string

Return type:

str

classmethod from_json(filename=None, string=None)[source]

Read a json formatted file or string, as written by to_json(). Exactly one of filename or string must be given.

Parameters:
  • filename (Optional[str]) – json filename

  • string (Optional[str]) – json formatted string: a list of objects with header and sequence keys

Returns:

New collection of the class this method is called on

Return type:

SequenceCollection

to_json(indent=None)[source]

json formatted string: a list of objects with header and sequence keys

Parameters:

indent (Optional[int]) – Indentation, passed to json.dumps()

Returns:

json formatted string

Return type:

str

abstractmethod pop(header)[source]

Remove a sequence from the collection and return it

Parameters:

header (str) – Header of the sequence to remove

Returns:

The removed sequence

Return type:

Sequence

Raises:

KeyError – If there is no sequence with this header

add(seq)[source]

Add a sequence to the collection (in place)

Parameters:

seq (Sequence) – Sequence to add

modify_inplace(mod_func)[source]

Change headers and/or sequences in place by calling mod_func on every sequence

Examples

>>> seqs = SequenceCollection.from_fasta(string='>a\nacgt\n>b\nggcc')
>>> seqs.modify_inplace(lambda header, sequence: (header.upper(), sequence.upper()))
>>> seqs.to_fasta()
'>A\nACGT\n>B\nGGCC'
Parameters:

mod_func (Callable[[str, str], tuple[str, str]]) – Function that takes a header and a sequence string, and returns a new header and sequence string

rename_inplace(rename_func)[source]

Rename all headers in place by calling rename_func on every header

Parameters:

rename_func (Callable[[str], str]) – Function that takes a header and returns a new header

class picea.SequenceReader(string=None, filename=None, filetype=None)[source]

Iterator over the sequences in a fasta or json formatted file or string. Exactly one of string or filename must be given. json input is a list of objects with header and sequence keys.

Examples

>>> fasta_string = '>1\nACGC\n>2\nTGTGTA\n'
>>> [seq.header for seq in SequenceReader(string=fasta_string, filetype='fasta')]
['1', '2']
Parameters:
  • string (str) – fasta or json formatted string

  • filename (str) – fasta or json filename

  • filetype (str) – "fasta" or "json"

Raises:

ValueError – If filetype is not supported

class picea.BatchSequenceReader(string=None, filename=None, filetype=None, batchsize=10)[source]

Iterator over batches of sequences in a fasta or json formatted file or string. Every batch is a SequenceCollection of batchsize sequences, except for the last batch, which can be smaller.

Examples

>>> fasta_string = '>1\nACGC\n>2\nTGTGTA\n>3\nAAGT\n>4\nCCA\n>5\nGGA\n'
>>> reader = BatchSequenceReader(string=fasta_string, filetype='fasta', batchsize=2)
>>> [batch.headers for batch in reader]
[['1', '2'], ['3', '4'], ['5']]
Parameters:
  • string (str) – fasta or json formatted string

  • filename (str) – fasta or json filename

  • filetype (str) – "fasta" or "json"

  • batchsize (int) – Number of sequences per batch