# Webhooks

> Get signed HTTPS callbacks when Jobs finish, balances run low and other events happen, and verify them.

Source: https://nodus-platform-site.pages.dev/docs/guides/webhooks/
Build revision: 211ad9f836655b1c3a2668c4693e442471f28614

A **WebhookEndpoint** sends an HTTPS request to your server each time an event you care about happens: a Job succeeds or fails, a balance runs low, a Budget crosses a threshold. You can have as many endpoints as you need, each with its own URL, event types and projects.

Every delivery is signed with the [Standard Webhooks](https://www.standardwebhooks.com) scheme, so you can check that it came from Nodus and was not changed in transit.

## Create an endpoint

endpoint.yaml

```yaml
apiVersion: nodus.dev/v1
kind: WebhookEndpoint
metadata:
  name: example-receiver
spec:
  url: https://hooks.example.com/nodus
  description: Job results for the example receiver
  eventTypes:
    - job.succeeded
    - job.failed
    - billing.low_balance
```

Terminal window

```console
$ nodus apply -f endpoint.yaml -o json | jq -r .status.signingSecret
whsec_...
```

The signing secret is in the response to the create, and nowhere else. Copy it into your receiver’s configuration now. If you lose it, [rotate it](https://nodus-platform-site.pages.dev/docs/guides/webhooks/#rotate-the-signing-secret) to get a new one.

Nodus sends a `webhookendpoint.ping` delivery right after the create, so you can see your receiver work before any real event.

|Field|Meaning|
|-|-|
|`url`|Where deliveries go. `https` only, with a public address. No redirects are followed.|
|`eventTypes`|Event types to send: an exact type (`job.succeeded`), a prefix (`job.*`, `billing.*`) or `*`. Defaults to `*`.|
|`projects`|Send project events only for these projects. Empty means every project. Org-level events always match.|
|`enabled`|`false` pauses deliveries. Pending deliveries to a paused endpoint are dropped.|
|`description`|A note for you.|

Every field can change after the create:

Terminal window

```console
$ nodus patch webhookendpoint ci --type merge -p '{"spec":{"eventTypes":["job.*"]}}'
```

## What Nodus sends

A delivery is a `POST` with a JSON body:

```json
{
  "type": "job.succeeded",
  "timestamp": "2026-10-01T12:00:00Z",
  "data": {
    "event": { "reason": "Succeeded", "message": "…", "count": 1 },
    "object": {
      "apiVersion": "nodus.dev/v1",
      "kind": "Job",
      "metadata": { "name": "train", "namespace": "default", "uid": "job_…" },
      "status": { "phase": "Succeeded" }
    }
  }
}
```

and three headers:

|Header|Meaning|
|-|-|
|`webhook-id`|Names this delivery. It stays the same across retries, so use it to ignore duplicates.|
|`webhook-timestamp`|When the attempt was signed, in Unix seconds.|
|`webhook-signature`|One or more signatures, separated by spaces, each like `v1,<base64>`.|

The event type is `<kind>.<reason>` in lowercase with underscores, such as `job.succeeded` or `agentrun.needs_resolution`. Billing events use the `billing.` prefix.

## Event types

|Type|When|
|-|-|
|`job.succeeded`, `job.failed`, `job.cancelled`|A Job reached a final phase. The same reasons exist for other run kinds, such as `agentrun.failed`.|
|`job.preempted`, `job.node_lost`, `job.restored`|A Job lost its machine, and was restored from its last checkpoint.|
|`job.checkpointed`, `job.checkpoint_failed`|A checkpoint was saved, or could not be.|
|`job.max_cost_reached`, `job.funding_lost`, `job.funding_restored`|A Job stopped or resumed because of its cost cap or your balance. The same reasons exist for Sandboxes and other kinds with a cost.|
|`job.output_sink_failed`|A Job’s output could not be loaded into its destination.|
|`job.gang_restarted`, `job.network_degraded`|A multi-node Job restarted or lost network quality.|
|`sandbox.started`, `sandbox.stopped`, `sandbox.stuck`, `sandbox.setup_failed`|A Sandbox changed state.|
|`agentrun.needs_resolution`|An AgentRun is waiting for you to resolve a step.|
|`billing.low_balance`|Your available balance fell below the warning level.|
|`billing.auto_recharge_succeeded`, `billing.auto_recharge_failed`|An auto-recharge added credits, or its charge failed.|
|`billing.arrears_posted`|A charge was larger than your balance, so new work waits until it is paid.|
|`billing.dispute_opened`, `billing.dispute_closed`|A payment was disputed, which pauses new work, or the dispute closed.|
|`billing.storage_quota_exceeded`, `billing.egress_quota_exceeded`|A quota was reached.|
|`budget.threshold`, `budget.exceeded`|A Budget crossed a threshold, or refused a new run.|
|`topup.succeeded`, `topup.failed`, `topup.refunded`|A top-up completed, its payment failed, or it was refunded.|
|`pool.node_offline`, `pool.forecast_shortfall`, `pool.action_proposed`|Events of your own capacity pools.|
|`webhookendpoint.ping`|Sent once when an endpoint is created.|
|`webhookendpoint.deliveries_failing`|Your receiver has failed every recent delivery.|

`nodus get events` shows the events of your org; each webhook type is the kind and reason of one of them.

## Verify the signature

Check every delivery before you trust it. The signature covers the exact bytes of the body, so verify the raw body before you parse it. Use a Standard Webhooks library for your language; it also rejects timestamps more than five minutes old, which stops a captured delivery from being replayed later.

main.go

```go
package main

import (
  "encoding/json"
  "fmt"
  "io"
  "log"
  "net/http"
  "os"
  "sync"
  "time"

  standardwebhooks "github.com/standard-webhooks/standard-webhooks/libraries/go"
)

// maxBody bounds a delivery; Nodus sends at most 64 KiB.
const maxBody = 64 << 10

// event is the body of a delivery: the event's type and the Event and object it is about.
type event struct {
  Type      string    `json:"type"`
  Timestamp time.Time `json:"timestamp"`
  Data      struct {
    Event  json.RawMessage `json:"event"`
    Object struct {
      Kind     string `json:"kind"`
      Metadata struct {
        Name      string `json:"name"`
        Namespace string `json:"namespace"`
      } `json:"metadata"`
    } `json:"object"`
  } `json:"data"`
}

// receiver verifies and handles deliveries.
type receiver struct {
  hook *standardwebhooks.Webhook
  out  io.Writer

  mu   sync.Mutex
  seen map[string]bool
}

func newReceiver(secret string, out io.Writer) (*receiver, error) {
  hook, err := standardwebhooks.NewWebhook(secret)
  if err != nil {
    return nil, fmt.Errorf("NODUS_WEBHOOK_SECRET: %w", err)
  }
  return &receiver{hook: hook, out: out, seen: map[string]bool{}}, nil
}

// ServeHTTP answers 2xx once the delivery is verified and recorded, which is all Nodus needs to stop retrying. A
// delivery that fails verification gets 400 and is never parsed: the signature covers the exact bytes received, so
// verify before you decode.
func (r *receiver) ServeHTTP(w http.ResponseWriter, req *http.Request) {
  body, err := io.ReadAll(http.MaxBytesReader(w, req.Body, maxBody))
  if err != nil {
    http.Error(w, "body too large", http.StatusRequestEntityTooLarge)
    return
  }
  // Verify checks the signature against every secret in the header and rejects timestamps more than five minutes
  // from now, so a captured delivery cannot be replayed later.
  if err := r.hook.Verify(body, req.Header); err != nil {
    http.Error(w, "invalid signature", http.StatusBadRequest)
    return
  }
  // A delivery can arrive more than once (a retry after a slow answer): webhook-id names the delivery, so handle
  // each id once and still answer 200.
  if r.firstTime(req.Header.Get(standardwebhooks.HeaderWebhookID)) {
    var e event
    if err := json.Unmarshal(body, &e); err != nil {
      http.Error(w, "malformed event", http.StatusBadRequest)
      return
    }
    r.handle(e)
  }
  w.WriteHeader(http.StatusOK)
}

func (r *receiver) firstTime(id string) bool {
  r.mu.Lock()
  defer r.mu.Unlock()
  if r.seen[id] {
    return false
  }
  r.seen[id] = true
  return true
}

// handle does the work of an event. Do slow work after answering (queue it); Nodus waits 15 seconds for a response.
func (r *receiver) handle(e event) {
  obj := e.Data.Object
  fmt.Fprintf(r.out, "%s %s/%s\n", e.Type, obj.Kind, obj.Metadata.Name)
}

func main() {
  rcv, err := newReceiver(os.Getenv("NODUS_WEBHOOK_SECRET"), os.Stdout)
  if err != nil {
    log.Fatal(err)
  }
  srv := &http.Server{Addr: ":8080", Handler: rcv, ReadHeaderTimeout: 5 * time.Second}
  log.Fatal(srv.ListenAndServe())
}
```

This receiver runs as is: `NODUS_WEBHOOK_SECRET=whsec_... go run ./examples/webhooks/receiver`. The same check in Python:

```python
from standardwebhooks.webhooks import Webhook

wh = Webhook(secret)  # the whsec_ secret
event = wh.verify(request_body, request_headers)  # raises if the signature or timestamp is wrong
```

If you cannot use a library, compute `base64(HMAC-SHA256(key, id + "." + timestamp + "." + body))`, where `key` is the secret after the `whsec_` prefix, base64-decoded, and compare it in constant time with each `v1,` signature in the header.

Caution

Answer with a `2xx` status as soon as you have verified and recorded the delivery, and do the slow work after. Nodus waits 15 seconds for a response, and treats a timeout like a failure.

## Retries and delivery order

A `2xx` response is success. Anything else, a timeout or a connection error is retried with exponential backoff, starting at 30 seconds and doubling up to 8 hours between attempts, for 3 days after the event. After that the delivery is marked failed. A URL that Nodus is not allowed to connect to, such as a private address, fails at once.

Deliveries can arrive more than once and out of order. Use `webhook-id` to ignore a duplicate, and the `timestamp` in the body or the object’s `status` to see which state is newer.

## The delivery log

Nodus keeps the log of an endpoint’s deliveries for 7 days:

Terminal window

```console
$ curl -s "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries?limit=20" \
    -H "Authorization: Bearer $NODUS_API_KEY"
```

Each entry shows the event type, the number of attempts, the receiver’s last response code, how long it took, and a `phase`: `Pending` while retries remain, then `Succeeded` or `Failed`. A delivery to an endpoint that was deleted or disabled is `Abandoned`. The response includes `continue` when there is more; pass it back as `after`.

To send a delivery again, replay it. The replay has the same body and a new `webhook-id`:

Terminal window

```console
$ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries/$ID/replays" \
    -H "Authorization: Bearer $NODUS_API_KEY"
```

## Rotate the signing secret

Rotate when a secret may have leaked, or when you lost it:

Terminal window

```console
$ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/rotations" \
    -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \
    -H "Idempotency-Key: $(uuidgen)" -d '{"previousSecretTTL": "24h"}' | jq -r .status.signingSecret
```

The response holds the new secret, once. For `previousSecretTTL` (24 hours by default, at most 7 days) every delivery carries a signature for the new secret and one for the old, so a receiver using either verifies it. Switch your receiver to the new secret, and the old one stops signing when the time is up. `status.secretRotation` shows when.

## Check an endpoint’s health

Terminal window

```console
$ nodus get webhookendpoint ci -o yaml
```

`status` shows the latest delivery (`lastDelivery`), the deliveries of the last 24 hours that succeeded or failed (`deliveries24h`), and, when the last three deliveries in a row failed, `failingSince` with the `Healthy` condition set to `False` and reason `DeliveriesFailing`. Nodus also records a `webhookendpoint.deliveries_failing` event when that happens, which you can send to another endpoint.

Note

Webhooks never carry provider names, machine ids or other internal details. Their bodies hold what the API would show you for the same object.
