nodus.envs
View MarkdownHelpers for building your own Environment and datasets: nodus.envs.split.
split keeps train and test disjoint under a canonical identity, the rule every catalog Environment follows:
two items with the same identity (the same question, the same graph up to relabeling, the same expression up to
commutativity) always land in the same split, so the held-out score measures generalisation, not recall.
train, test = nodus.envs.split(items, identity=lambda x: x["question"], test_size=0.2, seed=7)canonical
Section titled “canonical”canonical(value: Any) -> strA stable identity for any JSON-like value: sha256 of its sorted, compact JSON (strings are whitespace- and case-normalised, so trivially reformatted duplicates collide).
split(items: Iterable[T], identity: Callable[[T], Any] | None = None, test_size: float | int = 0.2, seed: int = 0) -> tuple[list[T], list[T]]Seeded train and test lists with no identity in both; each keeps the input order.
test_size is a fraction of the identity groups (0 < f < 1) or a number of groups. Groups, not items, are
assigned, so every duplicate of a held-out item is held out too.