betula has bindings for Python, TypeScript, and Rust. Each has one type per schema, named after the schema’s title, and a parse function that rejects exactly what the schema rejects. All three are tested against every example document on the schema pages.
Install¶
pip install betula-schemabetula-schema on PyPI. Requires Python 3.11+.
Pydantic v2 models, imported as betula_schema. parse and parse_json validate in Pydantic’s
strict mode, which is what makes them agree with the schema: calling Model.model_validate
directly would coerce, e.g., "1" into an int.
npm install betula-schemabetula-schema on npm. An ES module; its one runtime dependency is ajv.
Generated types, validated at runtime by ajv against the same schemas. parse(kind, data) is typed
by the kind name, so parse("Tree", data) returns a Tree.
cargo add betula-schemabetula-schema on crates.io, with API docs on docs.rs.
Serde types generated by typify, used as betula_schema::.... typify doesn’t enforce every JSON
Schema keyword, so parse and parse_str validate against the schema before deserializing.
Serializing and re-parsing gives an equal value, though empty optional arrays are omitted on output.
Parse a document¶
Parsing a Tree from a file. Every schema page shows the same code for its own type under Usage.
from betula_schema import Tree, parse_json
with open("tree.json") as f:
tree = parse_json(Tree, f.read())
print(tree)
import { readFileSync } from "node:fs";
import { parseJson, type Tree } from "betula-schema";
const tree: Tree = parseJson("Tree", readFileSync("tree.json", "utf8"));
console.log(tree);
fn main() -> Result<(), Box<dyn std::error::Error>> {
let text = std::fs::read_to_string("tree.json")?;
let tree: betula_schema::Tree = betula_schema::parse_str(&text)?;
println!("{tree:?}");
Ok(())
}
API tour¶
Each binding’s examples/usage.* file, included verbatim: union types, validation errors, strict
parsing, and serialization. CI runs these files, checks their assertions, and fails if a public
function or type isn’t used in them, so they cover the whole hand-written API.
"""Parsing betula JSON with the Python bindings. Run in CI; the asserts are checked."""
import json
from betula_schema import (
Alignment,
RnaSequence,
Sequence,
Tree,
UnwrappedAlignment,
ValidationError,
parse,
parse_json,
)
# Parse JSON text into a model. Every schema title is a class in betula_schema.
tree = parse_json(
Tree,
'{"name": "root", "length": 0, "children": [{"name": "A", "length": 0.1, "children": []}]}',
)
assert [child.name for child in tree.children] == ["A"]
# Or validate data you've already decoded. Parsing is strict: "0.1" is a
# string, not a number, so this is rejected rather than coerced.
try:
parse(Tree, {"name": "A", "length": "0.1", "children": []})
except ValidationError as err:
assert err.error_count() == 1
else:
raise AssertionError("expected a ValidationError")
# Union kinds (Sequence, Alignment) are RootModels; the variant that matched
# is `.root`. Constrained strings (Identifier, the alphabets) are RootModels
# too, so their plain value is also `.root`.
seq = parse(Sequence, {"type": "rna-sequence", "identifier": "s1", "sequence": "ACGU"})
match seq.root:
case RnaSequence(identifier=identifier, sequence=letters):
assert (identifier.root, letters.root) == ("s1", "ACGU")
case other:
raise AssertionError(f"unexpected variant {other!r}")
# A bare-array alignment matches UnwrappedAlignment, whose list is again `.root`.
alignment = parse(
Alignment,
[{"identifier": "s1", "sequence": "AC-GT"}, {"identifier": "s2", "sequence": "ACTGT"}],
)
assert isinstance(alignment.root, UnwrappedAlignment)
assert [s.root.identifier.root for s in alignment.root.root] == ["s1", "s2"]
# ValidationError lists every problem, with its location.
try:
parse_json(Sequence, '{"type": "dna-sequence", "identifier": "s2", "sequence": "ACGU"}')
except ValidationError as err:
assert any("pattern" in error["msg"] for error in err.errors())
else:
raise AssertionError("expected a ValidationError")
# Serialize with exclude_unset so optional fields you never set stay out.
assert json.loads(tree.model_dump_json(exclude_unset=True)) == {
"name": "root",
"length": 0.0,
"children": [{"name": "A", "length": 0.1, "children": []}],
}
print("python usage example ok")
The docstrings on parse and parse_json (help(betula_schema.parse)) carry examples that run as doctests.
// Parsing betula JSON with the TypeScript bindings. Compiled and run in CI;
// the asserts are checked.
import assert from "node:assert/strict";
import { BetulaValidationError, is, parse, parseJson, type Kind, type Sequence, type Tree } from "betula-schema";
// Parse JSON text. The kind name picks both the schema and the return type.
const tree: Tree = parseJson(
"Tree",
'{"name": "root", "length": 0, "children": [{"name": "A", "length": 0.1, "children": []}]}',
);
assert.deepEqual(tree.children.map((child) => child.name), ["A"]);
// Or validate data you've already decoded. parse() returns its input,
// typed; it doesn't copy or transform it.
const decoded: unknown = JSON.parse('{"identifier": "s1", "sequence": "ACGT"}');
const untyped = parse("Sequence", decoded);
assert.equal(untyped, decoded);
// Union kinds narrow on their `type` discriminator. UntypedSequence has no
// `type` field, so check for it first.
function describe(seq: Sequence): string {
if (!("type" in seq)) return `${seq.identifier}: unknown alphabet`;
switch (seq.type) {
case "dna-sequence":
return `${seq.identifier}: DNA`;
case "rna-sequence":
return `${seq.identifier}: RNA`;
case "protein-sequence":
return `${seq.identifier}: protein`;
}
}
assert.equal(describe(parse("Sequence", { type: "rna-sequence", identifier: "s2", sequence: "ACGU" })), "s2: RNA");
assert.equal(describe(untyped), "s1: unknown alphabet");
// is() is a type guard for branching instead of throwing.
const maybeTree: unknown = { name: "A", length: -1, children: [] };
assert.equal(is("Tree", maybeTree), false);
// Invalid data throws BetulaValidationError, carrying the kind and ajv's errors.
assert.throws(
() => parseJson("Sequence", '{"type": "dna-sequence", "identifier": "s3", "sequence": "ACGU"}'),
(err: unknown) => err instanceof BetulaValidationError && err.kind === "Sequence" && err.errors.length > 0,
);
// Kind is the union of every schema title, so generic helpers stay typed.
function count<K extends Kind>(kind: K, items: unknown[]): number {
return items.filter((item) => is(kind, item)).length;
}
assert.equal(count("Tree", [tree, maybeTree]), 1);
console.log("typescript usage example ok");
Signatures and TSDoc comments ship in the package’s .d.ts files, so they show up in your editor.
//! Parsing betula JSON with the Rust bindings. Run in CI; the asserts are checked.
use betula_schema::{Annotation, Sequence, Strand, Tree};
use serde_json::json;
fn main() -> Result<(), betula_schema::Error> {
// Parse JSON text. The type parameter picks both the schema and the result.
let tree: Tree = betula_schema::parse_str(
r#"{"name": "root", "length": 0, "children": [{"name": "A", "length": 0.1, "children": []}]}"#,
)?;
assert_eq!(tree.children[0].name, "A");
// Or validate a serde_json::Value you've already decoded.
let seq: Sequence = betula_schema::parse(json!({"type": "rna-sequence", "identifier": "s1", "sequence": "ACGU"}))?;
// Union kinds are enums. Constrained strings (Identifier, the alphabets)
// are newtypes that Deref to String.
match &seq {
Sequence::RnaSequence(rna) => assert_eq!((rna.identifier.as_str(), rna.sequence.as_str()), ("s1", "ACGU")),
other => panic!("unexpected variant {other:?}"),
}
// GFF3 strand is a plain enum.
let gene: Annotation = betula_schema::parse(json!({
"ID": "gene1", "seqid": "chr1", "source": "example", "interval_type": "gene",
"start": 1, "end": 900, "score": ".", "strand": "-", "phase": ".",
"attributes": {}, "children": []
}))?;
assert_eq!(gene.strand, Strand::Reverse);
// Invalid data is Error::Invalid, with one message per problem. parse
// checks the schema before deserializing, which matters: the generated
// types alone don't enforce every keyword.
match betula_schema::parse_str::<Sequence>(r#"{"type": "dna-sequence", "identifier": "s2", "sequence": "ACGU"}"#) {
Err(betula_schema::Error::Invalid { kind, errors }) => assert!(kind == "Sequence" && !errors.is_empty()),
other => panic!("expected Error::Invalid, got {other:?}"),
}
// validate() checks a Value without deserializing it.
assert!(betula_schema::validate::<Tree>(&json!({"name": "A", "length": -1, "children": []})).is_err());
// Serialize with serde as usual.
let text = serde_json::to_string(&tree).map_err(betula_schema::Error::Json)?;
assert_eq!(betula_schema::parse_str::<Tree>(&text)?, tree);
// Every generated type implements Kind, so generic code can take any of
// them; KINDS lists the names.
fn count_valid<K: betula_schema::Kind>(values: &[serde_json::Value]) -> usize {
values.iter().filter(|v| betula_schema::validate::<K>(v).is_ok()).count()
}
let candidates = [json!({"name": "A", "length": 1, "children": []}), json!({"name": "B"})];
assert_eq!(count_valid::<Tree>(&candidates), 1);
assert!(betula_schema::KINDS.contains(&<Tree as betula_schema::Kind>::NAME));
println!("rust usage example ok");
Ok(())
}
The API docs are on docs.rs; their examples run as doctests.
R¶
There’s no JSON-Schema-to-R code generator comparable to the ones above. From R, read with
jsonlite and validate with the jsonvalidate package against the self-contained bundle from
python3 scripts/bundle_schema.py (the individual schema files $ref each other by URIs that
don’t resolve over the network).