Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

ProteinAlphabet

https://schemas.wur.nl/betula/core/protein-alphabet/0.6.0/schema.json

IUPAC amino acid codes: the standard 20 (A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y) plus the ambiguity/special codes B (Asx), Z (Glx), X (any), J (Leu/Ile), U (selenocysteine), O (pyrrolysine) -- together these cover all 26 letters, so this pattern is effectively any letter. Also allows ‘*’ for a stop codon, and ‘-’/‘.’ as alignment gap/missing-data characters, since a Sequence doubles as an alignment row. Case-insensitive.

string (minLength 1, pattern ^[A-Za-z*.-]*$)

Valid examples

ambiguity-and-stop.json
"MKBZXJUO*-"

Invalid examples

digits.json
"MK123"
Schema source
protein-alphabet.schema.json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://schemas.wur.nl/betula/core/protein-alphabet/0.6.0/schema.json",
  "title": "ProteinAlphabet",
  "description": "IUPAC amino acid codes: the standard 20 (A,C,D,E,F,G,H,I,K,L,M,N,P,Q,R,S,T,V,W,Y) plus the ambiguity/special codes B (Asx), Z (Glx), X (any), J (Leu/Ile), U (selenocysteine), O (pyrrolysine) -- together these cover all 26 letters, so this pattern is effectively any letter. Also allows '*' for a stop codon, and '-'/'.' as alignment gap/missing-data characters, since a Sequence doubles as an alignment row. Case-insensitive.",
  "type": "string",
  "minLength": 1,
  "pattern": "^[A-Za-z*.-]*$"
}

Usage

Read a JSON document and parse it as a ProteinAlphabet. CI runs this exact code against the first valid example above; see Getting started to install the bindings.

Python
TypeScript
Rust
protein-alphabet.py
from betula_schema import ProteinAlphabet, parse_json

with open("protein-alphabet.json") as f:
    protein_alphabet = parse_json(ProteinAlphabet, f.read())
# ProteinAlphabet is a RootModel; protein_alphabet.root is the plain value
print(protein_alphabet)