https://schemas.wur.nl/betula/sequence/0.6.0/schema.json
A single biological sequence record, shared by picea, react-bio-viz, and acacia (FASTA-derived). ‘type’ is an optional alphabet discriminator: every current producer omits it and emits the plain UntypedSequence shape, which remains valid. Set it when the alphabet is known so consumers get both a proper discriminated union and IUPAC alphabet validation on ‘sequence’ (core/dna-alphabet, core/rna-alphabet, core/protein-alphabet) instead of an unconstrained string.
One of:
Definitions¶
DnaSequence¶
Field | Type | Required | Description |
|---|---|---|---|
|
| required | |
| required | ||
| required |
No properties beyond those listed are allowed.
RnaSequence¶
Field | Type | Required | Description |
|---|---|---|---|
|
| required | |
| required | ||
| required |
No properties beyond those listed are allowed.
ProteinSequence¶
Field | Type | Required | Description |
|---|---|---|---|
|
| required | |
| required | ||
| required |
No properties beyond those listed are allowed.
UntypedSequence¶
The shape every current producer (picea, react-bio-viz, acacia) actually emits: no alphabet discriminator.
Field | Type | Required | Description |
|---|---|---|---|
| required | ||
|
| required |
No properties beyond those listed are allowed.
Valid examples¶
{
"type": "dna-sequence",
"identifier": "seq6",
"sequence": "ACGT--RYSWKMBDHVN..acgtn"
}
{
"type": "dna-sequence",
"identifier": "seq1",
"sequence": "ACGTACGTNN"
}
{
"type": "protein-sequence",
"identifier": "seq7",
"sequence": "MKBZXJUO*-"
}
{
"type": "protein-sequence",
"identifier": "seq2",
"sequence": "MSTNPKPQRKTKRNTNRRPQDVKFPGGGQIVGGVYLLPRRGPRLGVRATRKTSERSQPR"
}
{
"type": "rna-sequence",
"identifier": "seq5",
"sequence": "ACGURYSWKMBDHVN"
}
{
"identifier": "seq3",
"sequence": "ACGTACGT"
}
Invalid examples¶
{
"type": "xyz-sequence",
"identifier": "seq4",
"sequence": "ACGT"
}
{
"type": "dna-sequence",
"identifier": "seq10",
"sequence": "ACGT123"
}
{
"type": "dna-sequence",
"identifier": "seq8",
"sequence": "ACGU"
}
{
"header": "seq4",
"sequence": "ACGT"
}
{
"identifier": "seq3"
}
{
"type": "protein-sequence",
"identifier": "seq11",
"sequence": "MK123"
}
{
"type": "rna-sequence",
"identifier": "seq9",
"sequence": "ACGT"
}
Schema source
{
"$schema": "https://json-schema.org/draft/2020-12/schema",
"$id": "https://schemas.wur.nl/betula/sequence/0.6.0/schema.json",
"title": "Sequence",
"description": "A single biological sequence record, shared by picea, react-bio-viz, and acacia (FASTA-derived). 'type' is an optional alphabet discriminator: every current producer omits it and emits the plain UntypedSequence shape, which remains valid. Set it when the alphabet is known so consumers get both a proper discriminated union and IUPAC alphabet validation on 'sequence' (core/dna-alphabet, core/rna-alphabet, core/protein-alphabet) instead of an unconstrained string.",
"oneOf": [
{ "$ref": "#/$defs/DnaSequence" },
{ "$ref": "#/$defs/RnaSequence" },
{ "$ref": "#/$defs/ProteinSequence" },
{ "$ref": "#/$defs/UntypedSequence" }
],
"$defs": {
"DnaSequence": {
"type": "object",
"properties": {
"type": { "const": "dna-sequence" },
"identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
"sequence": { "$ref": "https://schemas.wur.nl/betula/core/dna-alphabet/0.6.0/schema.json" }
},
"required": ["type", "identifier", "sequence"],
"additionalProperties": false
},
"RnaSequence": {
"type": "object",
"properties": {
"type": { "const": "rna-sequence" },
"identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
"sequence": { "$ref": "https://schemas.wur.nl/betula/core/rna-alphabet/0.6.0/schema.json" }
},
"required": ["type", "identifier", "sequence"],
"additionalProperties": false
},
"ProteinSequence": {
"type": "object",
"properties": {
"type": { "const": "protein-sequence" },
"identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
"sequence": { "$ref": "https://schemas.wur.nl/betula/core/protein-alphabet/0.6.0/schema.json" }
},
"required": ["type", "identifier", "sequence"],
"additionalProperties": false
},
"UntypedSequence": {
"type": "object",
"description": "The shape every current producer (picea, react-bio-viz, acacia) actually emits: no alphabet discriminator.",
"properties": {
"identifier": { "$ref": "https://schemas.wur.nl/betula/core/identifier/0.6.0/schema.json" },
"sequence": { "type": "string", "minLength": 1 }
},
"required": ["identifier", "sequence"],
"additionalProperties": false
}
}
}
Usage¶
Read a JSON document and parse it as a Sequence. CI runs this exact code against the first valid example above; see Getting started to install the bindings.
from betula_schema import Sequence, parse_json
with open("sequence.json") as f:
sequence = parse_json(Sequence, f.read())
# sequence.root is the matched variant: DnaSequence, RnaSequence, ProteinSequence, or UntypedSequence
print(sequence)
import { readFileSync } from "node:fs";
import { parseJson, type Sequence } from "betula-schema";
const sequence: Sequence = parseJson("Sequence", readFileSync("sequence.json", "utf8"));
// Sequence is a union type: DnaSequence | RnaSequence | ProteinSequence | UntypedSequence
console.log(sequence);
fn main() -> Result<(), Box<dyn std::error::Error>> {
let text = std::fs::read_to_string("sequence.json")?;
let sequence: betula_schema::Sequence = betula_schema::parse_str(&text)?;
// Sequence is an enum with one variant per alternative, e.g. Sequence::DnaSequence(_)
println!("{sequence:?}");
Ok(())
}