Skip to content

nodus.integrations.transformers

View Markdown

NodusCallback for the Hugging Face Trainer and every TRL trainer built on it (ADR-103).

from nodus.integrations.transformers import NodusCallback
trainer = SFTTrainer(..., callbacks=[NodusCallback()])
trainer.train(resume_from_checkpoint=nodus.integrations.transformers.last_checkpoint())

Point output_dir at nodus.state_dir() so the Trainer’s checkpoint-<step> folders are the recovery state.

class NodusCallback(phase: str = 'Training') -> None

Metrics, progress, phases and the checkpoint handshake for a transformers.Trainer.

on_log(args: Any, state: Any, control: Any, logs: dict[str, Any] | None = None, **kwargs: Any) -> None
on_save(args: Any, state: Any, control: Any, **kwargs: Any) -> None
on_step_end(args: Any, state: Any, control: Any, **kwargs: Any) -> Any
on_train_begin(args: Any, state: Any, control: Any, **kwargs: Any) -> None
last_checkpoint(directory: str | os.PathLike[str] | None = None) -> str | None

The newest complete checkpoint-<step> under directory (default the state dir), for resume_from_checkpoint.

A folder counts only when the Trainer finished writing it (trainer_state.json is present), so an attempt stopped mid-save resumes from the checkpoint before it.