# nodus.envs

> Helpers for building your own Environment and datasets: `nodus.envs.split`.

Source: https://nodus-platform-site.pages.dev/docs/reference/python/nodus-envs/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

<!-- Generated by scripts/gen-reference.mjs from the SDK docstrings (griffe). Do not edit: run make gen. -->

Helpers for building your own Environment and datasets: `nodus.envs.split`.

`split` keeps train and test disjoint under a canonical identity, the rule every catalog Environment follows: two items with the same identity (the same question, the same graph up to relabeling, the same expression up to commutativity) always land in the same split, so the held-out score measures generalisation, not recall.

```plaintext
train, test = nodus.envs.split(items, identity=lambda x: x["question"], test_size=0.2, seed=7)
```

## `canonical`

```python
canonical(value: Any) -> str
```

A stable identity for any JSON-like value: sha256 of its sorted, compact JSON (strings are whitespace- and case-normalised, so trivially reformatted duplicates collide).

## `split`

```python
split(items: Iterable[T], identity: Callable[[T], Any] | None = None, test_size: float | int = 0.2, seed: int = 0) -> tuple[list[T], list[T]]
```

Seeded train and test lists with no identity in both; each keeps the input order.

`test_size` is a fraction of the identity groups (0 < f < 1) or a number of groups. Groups, not items, are assigned, so every duplicate of a held-out item is held out too.
