# Functions and classes

> Define Functions with @app.function, call them with .remote, .map and .spawn, keep state in @app.cls classes, and deploy Apps.

Source: https://nodus-platform-site.pages.dev/docs/guides/python/functions/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A Function is a Python function that runs on Nodus workers. Workers start when calls arrive, stay warm for `scaledown_window` and scale between `min_workers` and `max_workers`. Every call is a FunctionCall object, so you can look it up, wait for it later or cancel it.

## Define a Function

```python
import nodus

app = nodus.App("finetune")
image = nodus.Image.debian_slim().pip_install("torch", "transformers==4.57.6")

@app.function(gpu="H100", image=image, timeout="6h", checkpoint="/nodus/state")
def train(lr: float) -> dict:
    ...
```

|Argument|What it sets|
|-|-|
|`gpu`|`"H100"`, `"H100:2"`, `"A100-80GB"`, `"H100!"` (exact variant), a list of alternatives, or `nodus.GPU(...)`|
|`cpu`, `memory`, `ephemeral_disk`|Floors per worker; bare numbers are vCPUs and MiB|
|`image`|A `nodus.Image`; defaults to `nodus/python` at your Python version|
|`secrets`, `volumes`, `env`, `network`|`[nodus.Secret]`, `{"/path": nodus.Volume}`, a dict, `nodus.Egress.deny()`/`.allow(...)`/`.open()`|
|`timeout`, `retries`|Per-call limit (`"6h"` or seconds); retries for exceptions (`int` or `nodus.Retries(...)`)|
|`checkpoint`|A path (or `True` for `/nodus/state`) saved and restored when a worker is lost|
|`interruptible`, `region`, `profile`|Allow interruptible capacity; region classes such as `["us", "eu"]`; `Balanced`, `Cost` or `Speed`|
|`max_cost`|Not available for Functions yet: deploying with it is refused (Jobs and Sandboxes take it)|
|`min_workers`, `max_workers`, `scaledown_window`, `target_concurrency`|Worker pool sizing (`min_containers` and `max_containers` also work)|
|`name`|The member name; the Function object is `<app>-<name>`|

Decorating does not contact Nodus, so importing the file has no side effects. The App’s code (the directory of the file, minus what `.gitignore` and `.nodusignore` exclude) is uploaded once per content hash when the App runs.

## Call it

```python
with app.run():                         # or: nodus run app.py
    result = train.remote(3e-4)         # one call; blocks and returns the result
    call = train.spawn(1e-4)            # starts a call and returns a handle
    print(call.get(timeout=3600))
    for y in train.map([1e-4, 3e-4]):   # one call per input, results in input order
        print(y)
    print(train.estimate(3e-4))         # expected cost, start time and hold, without running
```

* `.map(*iterables, order_outputs=True, return_exceptions=False)` creates calls in batches of 1,000. With `return_exceptions=True` a failed input yields its exception instead of stopping the loop.
* `.starmap(pairs)` spreads each tuple into arguments; `.for_each(xs)` runs and discards the results.
* `nodus.FunctionCall.from_name(name)` finds a spawned call again, from any process.
* Arguments and results are serialized with cloudpickle. Values above 64 KiB travel as uploaded blobs.

An exception raised in the Function is raised again in your process as its own type when it can be imported there; otherwise you get `nodus.errors.RemoteError`. Either way the remote traceback is attached.

## Deploy and look up

Terminal window

```console
nodus deploy app.py
```

```python
train = nodus.Function.from_name("finetune", "train")
print(train.remote(3e-4))
```

A deploy updates the App in place: changed Functions roll their workers after in-flight calls finish, and Functions you removed from the file are deleted.

## Classes

`@app.cls` turns a class into one Function. `@nodus.enter()` methods run once per worker, before its first call, so expensive setup such as loading a model happens once; `@nodus.exit()` methods run when the worker drains. Methods marked `@nodus.method()` get `.remote()`, `.map()` and `.spawn()`.

examples/python/classes/app.py

```python
"""A class whose model loads once per worker, then serves many calls.

Run it with `nodus run examples/python/classes/app.py`.
"""

import nodus

app = nodus.App("classes")


@app.cls(cpu=2, memory="4Gi", scaledown_window="5m", max_cost=1)
class Greeter:
    @nodus.enter()
    def load(self) -> None:
        # Runs once when a worker starts, before its first call: load weights or open connections here.
        self.greeting = "hello"

    @nodus.method()
    def greet(self, name: str) -> str:
        return f"{self.greeting}, {name}"

    @nodus.exit()
    def close(self) -> None:
        self.greeting = ""


@app.local_entrypoint()
def main() -> None:
    greeter = Greeter()
    print(greeter.greet.remote("Ada"))
    print(list(greeter.greet.map(["Grace", "Linus"])))
```

Classes take no constructor arguments; configure them in the `enter` hook.

## Clustered Functions (Beta)

Beta

Clustered Functions need the distributed training beta for your org. Checkpointed resume of a gang is not yet qualified.

`@nodus.clustered(size=N)` below `@app.function` runs each `.remote()` or `.spawn()` as a gang of `N` nodes. Every member runs the function; rank 0’s return value is the result. `nodus.cluster.info()` tells each member its rank, the gang size, the member addresses and the rendezvous address, so `torch.distributed` initializes from the environment.

examples/functions/clustered/app.py

```python
"""A clustered Function: each call runs as one gang of two nodes, and rank 0's return value is the result (Beta).

Run it with `nodus run examples/functions/clustered/app.py`.
"""

import nodus

app = nodus.App("clustered")
image = nodus.Image.from_registry("nodus/pytorch:2.8-cuda12.8")


@app.function(gpu="H100", image=image, timeout="30m", max_cost=1)
@nodus.clustered(size=2)
def whoami() -> dict:
    info = nodus.cluster.info()  # rank, size, member addresses, rendezvous address and epoch
    print(f"rank {info.rank} of {info.size}; rendezvous {info.master_addr}:{info.master_port}")
    return {"rank": info.rank, "size": info.size}


@app.local_entrypoint()
def main() -> None:
    result = whoami.remote()
    assert result == {"rank": 0, "size": 2}, result
    print(result)
```

`network` (`Colocated`, `Regional`, `Global`) and `transport` (`Direct`, `Auto`) choose where members may be placed. Clustered Functions keep no warm workers, and `.map()` on them raises `nodus.errors.Unsupported`.
