Skip to content

nodus.envs

View Markdown

Helpers for building your own Environment and datasets: nodus.envs.split.

split keeps train and test disjoint under a canonical identity, the rule every catalog Environment follows: two items with the same identity (the same question, the same graph up to relabeling, the same expression up to commutativity) always land in the same split, so the held-out score measures generalisation, not recall.

train, test = nodus.envs.split(items, identity=lambda x: x["question"], test_size=0.2, seed=7)
canonical(value: Any) -> str

A stable identity for any JSON-like value: sha256 of its sorted, compact JSON (strings are whitespace- and case-normalised, so trivially reformatted duplicates collide).

split(items: Iterable[T], identity: Callable[[T], Any] | None = None, test_size: float | int = 0.2, seed: int = 0) -> tuple[list[T], list[T]]

Seeded train and test lists with no identity in both; each keeps the input order.

test_size is a fraction of the identity groups (0 < f < 1) or a number of groups. Groups, not items, are assigned, so every duplicate of a held-out item is held out too.