nodus.integrations.transformers
View MarkdownNodusCallback for the Hugging Face Trainer and every TRL trainer built on it (ADR-103).
from nodus.integrations.transformers import NodusCallback
trainer = SFTTrainer(..., callbacks=[NodusCallback()])trainer.train(resume_from_checkpoint=nodus.integrations.transformers.last_checkpoint())Point output_dir at nodus.state_dir() so the Trainer’s checkpoint-<step> folders are the recovery state.
NodusCallback
Section titled “NodusCallback”class NodusCallback(phase: str = 'Training') -> NoneMetrics, progress, phases and the checkpoint handshake for a transformers.Trainer.
NodusCallback.on_log
Section titled “NodusCallback.on_log”on_log(args: Any, state: Any, control: Any, logs: dict[str, Any] | None = None, **kwargs: Any) -> NoneNodusCallback.on_save
Section titled “NodusCallback.on_save”on_save(args: Any, state: Any, control: Any, **kwargs: Any) -> NoneNodusCallback.on_step_end
Section titled “NodusCallback.on_step_end”on_step_end(args: Any, state: Any, control: Any, **kwargs: Any) -> AnyNodusCallback.on_train_begin
Section titled “NodusCallback.on_train_begin”on_train_begin(args: Any, state: Any, control: Any, **kwargs: Any) -> Nonelast_checkpoint
Section titled “last_checkpoint”last_checkpoint(directory: str | os.PathLike[str] | None = None) -> str | NoneThe newest complete checkpoint-<step> under directory (default the state dir), for resume_from_checkpoint.
A folder counts only when the Trainer finished writing it (trainer_state.json is present), so an attempt
stopped mid-save resumes from the checkpoint before it.