Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Getting started

betula has bindings for Python, TypeScript, and Rust. Each has one type per schema, named after the schema’s title, and a parse function that rejects exactly what the schema rejects. All three are tested against every example document on the schema pages.

Install

Python
TypeScript
Rust
pip install betula-schema

betula-schema on PyPI. Requires Python 3.11+.

Pydantic v2 models, imported as betula_schema. parse and parse_json validate in Pydantic’s strict mode, which is what makes them agree with the schema: calling Model.model_validate directly would coerce, e.g., "1" into an int.

Parse a document

Parsing a Tree from a file. Every schema page shows the same code for its own type under Usage.

Python
TypeScript
Rust
tree.py
from betula_schema import Tree, parse_json

with open("tree.json") as f:
    tree = parse_json(Tree, f.read())
print(tree)

API tour

Each binding’s examples/usage.* file, included verbatim: union types, validation errors, strict parsing, and serialization. CI runs these files, checks their assertions, and fails if a public function or type isn’t used in them, so they cover the whole hand-written API.

Python
TypeScript
Rust
usage.py
"""Parsing betula JSON with the Python bindings. Run in CI; the asserts are checked."""

import json

from betula_schema import (
    Alignment,
    RnaSequence,
    Sequence,
    Tree,
    UnwrappedAlignment,
    ValidationError,
    parse,
    parse_json,
)

# Parse JSON text into a model. Every schema title is a class in betula_schema.
tree = parse_json(
    Tree,
    '{"name": "root", "length": 0, "children": [{"name": "A", "length": 0.1, "children": []}]}',
)
assert [child.name for child in tree.children] == ["A"]

# Or validate data you've already decoded. Parsing is strict: "0.1" is a
# string, not a number, so this is rejected rather than coerced.
try:
    parse(Tree, {"name": "A", "length": "0.1", "children": []})
except ValidationError as err:
    assert err.error_count() == 1
else:
    raise AssertionError("expected a ValidationError")

# Union kinds (Sequence, Alignment) are RootModels; the variant that matched
# is `.root`. Constrained strings (Identifier, the alphabets) are RootModels
# too, so their plain value is also `.root`.
seq = parse(Sequence, {"type": "rna-sequence", "identifier": "s1", "sequence": "ACGU"})
match seq.root:
    case RnaSequence(identifier=identifier, sequence=letters):
        assert (identifier.root, letters.root) == ("s1", "ACGU")
    case other:
        raise AssertionError(f"unexpected variant {other!r}")

# A bare-array alignment matches UnwrappedAlignment, whose list is again `.root`.
alignment = parse(
    Alignment,
    [{"identifier": "s1", "sequence": "AC-GT"}, {"identifier": "s2", "sequence": "ACTGT"}],
)
assert isinstance(alignment.root, UnwrappedAlignment)
assert [s.root.identifier.root for s in alignment.root.root] == ["s1", "s2"]

# ValidationError lists every problem, with its location.
try:
    parse_json(Sequence, '{"type": "dna-sequence", "identifier": "s2", "sequence": "ACGU"}')
except ValidationError as err:
    assert any("pattern" in error["msg"] for error in err.errors())
else:
    raise AssertionError("expected a ValidationError")

# Serialize with exclude_unset so optional fields you never set stay out.
assert json.loads(tree.model_dump_json(exclude_unset=True)) == {
    "name": "root",
    "length": 0.0,
    "children": [{"name": "A", "length": 0.1, "children": []}],
}
print("python usage example ok")

The docstrings on parse and parse_json (help(betula_schema.parse)) carry examples that run as doctests.

R

There’s no JSON-Schema-to-R code generator comparable to the ones above. From R, read with jsonlite and validate with the jsonvalidate package against the self-contained bundle from python3 scripts/bundle_schema.py (the individual schema files $ref each other by URIs that don’t resolve over the network).