Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Sequence

https://schemas.wur.nl/betula/sequence/0.6.0/schema.json

A single biological sequence record, shared by picea, react-bio-viz, and acacia (FASTA-derived). ‘type’ is an optional alphabet discriminator: every current producer omits it and emits the plain UntypedSequence shape, which remains valid. Set it when the alphabet is known so consumers get both a proper discriminated union and IUPAC alphabet validation on ‘sequence’ (core/dna-alphabet, core/rna-alphabet, core/protein-alphabet) instead of an unconstrained string.

One of:

Definitions

DnaSequence

Field

Type

Required

Description

type

"dna-sequence"

required

identifier

Identifier

required

sequence

DnaAlphabet

required

No properties beyond those listed are allowed.

RnaSequence

Field

Type

Required

Description

type

"rna-sequence"

required

identifier

Identifier

required

sequence

RnaAlphabet

required

No properties beyond those listed are allowed.

ProteinSequence

Field

Type

Required

Description

type

"protein-sequence"

required

identifier

Identifier

required

sequence

ProteinAlphabet

required

No properties beyond those listed are allowed.

UntypedSequence

The shape every current producer (picea, react-bio-viz, acacia) actually emits: no alphabet discriminator.

Field

Type

Required

Description

identifier

Identifier

required

sequence

string (minLength 1)

required

No properties beyond those listed are allowed.

Valid examples

dna-with-ambiguity-and-gaps.json
{
  "type": "dna-sequence",
  "identifier": "seq6",
  "sequence": "ACGT--RYSWKMBDHVN..acgtn"
}
dna.json
{
  "type": "dna-sequence",
  "identifier": "seq1",
  "sequence": "ACGTACGTNN"
}
protein-with-ambiguity.json
{
  "type": "protein-sequence",
  "identifier": "seq7",
  "sequence": "MKBZXJUO*-"
}
protein.json
{
  "type": "protein-sequence",
  "identifier": "seq2",
  "sequence": "MSTNPKPQRKTKRNTNRRPQDVKFPGGGQIVGGVYLLPRRGPRLGVRATRKTSERSQPR"
}
rna.json
{
  "type": "rna-sequence",
  "identifier": "seq5",
  "sequence": "ACGURYSWKMBDHVN"
}
untyped.json
{
  "identifier": "seq3",
  "sequence": "ACGTACGT"
}

Invalid examples

bad-type.json
{
  "type": "xyz-sequence",
  "identifier": "seq4",
  "sequence": "ACGT"
}
dna-bad-chars.json
{
  "type": "dna-sequence",
  "identifier": "seq10",
  "sequence": "ACGT123"
}
dna-contains-u.json
{
  "type": "dna-sequence",
  "identifier": "seq8",
  "sequence": "ACGU"
}
legacy-header-field.json
{
  "header": "seq4",
  "sequence": "ACGT"
}
missing-sequence.json
{
  "identifier": "seq3"
}
protein-bad-chars.json
{
  "type": "protein-sequence",
  "identifier": "seq11",
  "sequence": "MK123"
}
rna-contains-t.json
{
  "type": "rna-sequence",
  "identifier": "seq9",
  "sequence": "ACGT"
}
Schema source
sequence.schema.json
{
  "$schema": "https://json-schema.org/draft/2020-12/schema",
  "$id": "https://schemas.wur.nl/betula/sequence/0.6.0/schema.json",
  "title": "Sequence",
  "description": "A single biological sequence record, shared by picea, react-bio-viz, and acacia (FASTA-derived). 'type' is an optional alphabet discriminator: every current producer omits it and emits the plain UntypedSequence shape, which remains valid. Set it when the alphabet is known so consumers get both a proper discriminated union and IUPAC alphabet validation on 'sequence' (core/dna-alphabet, core/rna-alphabet, core/protein-alphabet) instead of an unconstrained string.",
  "oneOf": [
    { "$ref": "#/$defs/DnaSequence" },
    { "$ref": "#/$defs/RnaSequence" },
    { "$ref": "#/$defs/ProteinSequence" },
    { "$ref": "#/$defs/UntypedSequence" }
  ],
  "$defs": {
    "DnaSequence": {
      "type": "object",
      "properties": {
        "type": { "const": "dna-sequence" },
        "identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
        "sequence": { "$ref": "https://schemas.wur.nl/betula/core/dna-alphabet/0.6.0/schema.json" }
      },
      "required": ["type", "identifier", "sequence"],
      "additionalProperties": false
    },
    "RnaSequence": {
      "type": "object",
      "properties": {
        "type": { "const": "rna-sequence" },
        "identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
        "sequence": { "$ref": "https://schemas.wur.nl/betula/core/rna-alphabet/0.6.0/schema.json" }
      },
      "required": ["type", "identifier", "sequence"],
      "additionalProperties": false
    },
    "ProteinSequence": {
      "type": "object",
      "properties": {
        "type": { "const": "protein-sequence" },
        "identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
        "sequence": { "$ref": "https://schemas.wur.nl/betula/core/protein-alphabet/0.6.0/schema.json" }
      },
      "required": ["type", "identifier", "sequence"],
      "additionalProperties": false
    },
    "UntypedSequence": {
      "type": "object",
      "description": "The shape every current producer (picea, react-bio-viz, acacia) actually emits: no alphabet discriminator.",
      "properties": {
        "identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
        "sequence": { "type": "string", "minLength": 1 }
      },
      "required": ["identifier", "sequence"],
      "additionalProperties": false
    }
  }
}

Usage

Read a JSON document and parse it as a Sequence. CI runs this exact code against the first valid example above; see Getting started to install the bindings.

Python
TypeScript
Rust
sequence.py
from betula_schema import Sequence, parse_json

with open("sequence.json") as f:
    sequence = parse_json(Sequence, f.read())
# sequence.root is the matched variant: DnaSequence, RnaSequence, ProteinSequence, or UntypedSequence
print(sequence)