Webhooks
View MarkdownA WebhookEndpoint sends an HTTPS request to your server each time an event you care about happens: a Job succeeds or fails, a balance runs low, a Budget crosses a threshold. You can have as many endpoints as you need, each with its own URL, event types and projects.
Every delivery is signed with the Standard Webhooks scheme, so you can check that it came from Nodus and was not changed in transit.
Create an endpoint
Section titled “Create an endpoint”apiVersion: nodus.dev/v1kind: WebhookEndpointmetadata: name: example-receiverspec: url: https://hooks.example.com/nodus description: Job results for the example receiver eventTypes: - job.succeeded - job.failed - billing.low_balance$ nodus apply -f endpoint.yaml -o json | jq -r .status.signingSecretwhsec_...The signing secret is in the response to the create, and nowhere else. Copy it into your receiver’s configuration now. If you lose it, rotate it to get a new one.
Nodus sends a webhookendpoint.ping delivery right after the create, so you can see your receiver work before any
real event.
| Field | Meaning |
|---|---|
url |
Where deliveries go. https only, with a public address. No redirects are followed. |
eventTypes |
Event types to send: an exact type (job.succeeded), a prefix (job.*, billing.*) or *. Defaults to *. |
projects |
Send project events only for these projects. Empty means every project. Org-level events always match. |
enabled |
false pauses deliveries. Pending deliveries to a paused endpoint are dropped. |
description |
A note for you. |
Every field can change after the create:
$ nodus patch webhookendpoint ci --type merge -p '{"spec":{"eventTypes":["job.*"]}}'What Nodus sends
Section titled “What Nodus sends”A delivery is a POST with a JSON body:
{ "type": "job.succeeded", "timestamp": "2026-10-01T12:00:00Z", "data": { "event": { "reason": "Succeeded", "message": "…", "count": 1 }, "object": { "apiVersion": "nodus.dev/v1", "kind": "Job", "metadata": { "name": "train", "namespace": "default", "uid": "job_…" }, "status": { "phase": "Succeeded" } } }}and three headers:
| Header | Meaning |
|---|---|
webhook-id |
Names this delivery. It stays the same across retries, so use it to ignore duplicates. |
webhook-timestamp |
When the attempt was signed, in Unix seconds. |
webhook-signature |
One or more signatures, separated by spaces, each like v1,<base64>. |
The event type is <kind>.<reason> in lowercase with underscores, such as job.succeeded or
agentrun.needs_resolution. Billing events use the billing. prefix.
Event types
Section titled “Event types”| Type | When |
|---|---|
job.succeeded, job.failed, job.cancelled |
A Job reached a final phase. The same reasons exist for other run kinds, such as agentrun.failed. |
job.preempted, job.node_lost, job.restored |
A Job lost its machine, and was restored from its last checkpoint. |
job.checkpointed, job.checkpoint_failed |
A checkpoint was saved, or could not be. |
job.max_cost_reached, job.funding_lost, job.funding_restored |
A Job stopped or resumed because of its cost cap or your balance. The same reasons exist for Sandboxes and other kinds with a cost. |
job.output_sink_failed |
A Job’s output could not be loaded into its destination. |
job.gang_restarted, job.network_degraded |
A multi-node Job restarted or lost network quality. |
sandbox.started, sandbox.stopped, sandbox.stuck, sandbox.setup_failed |
A Sandbox changed state. |
agentrun.needs_resolution |
An AgentRun is waiting for you to resolve a step. |
billing.low_balance |
Your available balance fell below the warning level. |
billing.auto_recharge_succeeded, billing.auto_recharge_failed |
An auto-recharge added credits, or its charge failed. |
billing.arrears_posted |
A charge was larger than your balance, so new work waits until it is paid. |
billing.dispute_opened, billing.dispute_closed |
A payment was disputed, which pauses new work, or the dispute closed. |
billing.storage_quota_exceeded, billing.egress_quota_exceeded |
A quota was reached. |
budget.threshold, budget.exceeded |
A Budget crossed a threshold, or refused a new run. |
topup.succeeded, topup.failed, topup.refunded |
A top-up completed, its payment failed, or it was refunded. |
pool.node_offline, pool.forecast_shortfall, pool.action_proposed |
Events of your own capacity pools. |
webhookendpoint.ping |
Sent once when an endpoint is created. |
webhookendpoint.deliveries_failing |
Your receiver has failed every recent delivery. |
nodus get events shows the events of your org; each webhook type is the kind and reason of one of them.
Verify the signature
Section titled “Verify the signature”Check every delivery before you trust it. The signature covers the exact bytes of the body, so verify the raw body before you parse it. Use a Standard Webhooks library for your language; it also rejects timestamps more than five minutes old, which stops a captured delivery from being replayed later.
package main
import ( "encoding/json" "fmt" "io" "log" "net/http" "os" "sync" "time"
standardwebhooks "github.com/standard-webhooks/standard-webhooks/libraries/go")
// maxBody bounds a delivery; Nodus sends at most 64 KiB.const maxBody = 64 << 10
// event is the body of a delivery: the event's type and the Event and object it is about.type event struct { Type string `json:"type"` Timestamp time.Time `json:"timestamp"` Data struct { Event json.RawMessage `json:"event"` Object struct { Kind string `json:"kind"` Metadata struct { Name string `json:"name"` Namespace string `json:"namespace"` } `json:"metadata"` } `json:"object"` } `json:"data"`}
// receiver verifies and handles deliveries.type receiver struct { hook *standardwebhooks.Webhook out io.Writer
mu sync.Mutex seen map[string]bool}
func newReceiver(secret string, out io.Writer) (*receiver, error) { hook, err := standardwebhooks.NewWebhook(secret) if err != nil { return nil, fmt.Errorf("NODUS_WEBHOOK_SECRET: %w", err) } return &receiver{hook: hook, out: out, seen: map[string]bool{}}, nil}
// ServeHTTP answers 2xx once the delivery is verified and recorded, which is all Nodus needs to stop retrying. A// delivery that fails verification gets 400 and is never parsed: the signature covers the exact bytes received, so// verify before you decode.func (r *receiver) ServeHTTP(w http.ResponseWriter, req *http.Request) { body, err := io.ReadAll(http.MaxBytesReader(w, req.Body, maxBody)) if err != nil { http.Error(w, "body too large", http.StatusRequestEntityTooLarge) return } // Verify checks the signature against every secret in the header and rejects timestamps more than five minutes // from now, so a captured delivery cannot be replayed later. if err := r.hook.Verify(body, req.Header); err != nil { http.Error(w, "invalid signature", http.StatusBadRequest) return } // A delivery can arrive more than once (a retry after a slow answer): webhook-id names the delivery, so handle // each id once and still answer 200. if r.firstTime(req.Header.Get(standardwebhooks.HeaderWebhookID)) { var e event if err := json.Unmarshal(body, &e); err != nil { http.Error(w, "malformed event", http.StatusBadRequest) return } r.handle(e) } w.WriteHeader(http.StatusOK)}
func (r *receiver) firstTime(id string) bool { r.mu.Lock() defer r.mu.Unlock() if r.seen[id] { return false } r.seen[id] = true return true}
// handle does the work of an event. Do slow work after answering (queue it); Nodus waits 15 seconds for a response.func (r *receiver) handle(e event) { obj := e.Data.Object fmt.Fprintf(r.out, "%s %s/%s\n", e.Type, obj.Kind, obj.Metadata.Name)}
func main() { rcv, err := newReceiver(os.Getenv("NODUS_WEBHOOK_SECRET"), os.Stdout) if err != nil { log.Fatal(err) } srv := &http.Server{Addr: ":8080", Handler: rcv, ReadHeaderTimeout: 5 * time.Second} log.Fatal(srv.ListenAndServe())}This receiver runs as is: NODUS_WEBHOOK_SECRET=whsec_... go run ./examples/webhooks/receiver. The same check in
Python:
from standardwebhooks.webhooks import Webhook
wh = Webhook(secret) # the whsec_ secretevent = wh.verify(request_body, request_headers) # raises if the signature or timestamp is wrongIf you cannot use a library, compute base64(HMAC-SHA256(key, id + "." + timestamp + "." + body)), where key is
the secret after the whsec_ prefix, base64-decoded, and compare it in constant time with each v1, signature in
the header.
Retries and delivery order
Section titled “Retries and delivery order”A 2xx response is success. Anything else, a timeout or a connection error is retried with exponential backoff,
starting at 30 seconds and doubling up to 8 hours between attempts, for 3 days after the event. After that the
delivery is marked failed. A URL that Nodus is not allowed to connect to, such as a private address, fails at once.
Deliveries can arrive more than once and out of order. Use webhook-id to ignore a duplicate, and the timestamp
in the body or the object’s status to see which state is newer.
The delivery log
Section titled “The delivery log”Nodus keeps the log of an endpoint’s deliveries for 7 days:
$ curl -s "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries?limit=20" \ -H "Authorization: Bearer $NODUS_API_KEY"Each entry shows the event type, the number of attempts, the receiver’s last response code, how long it took, and
a phase: Pending while retries remain, then Succeeded or Failed. A delivery to an endpoint that was deleted
or disabled is Abandoned. The response includes continue when there is more; pass it back as after.
To send a delivery again, replay it. The replay has the same body and a new webhook-id:
$ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries/$ID/replays" \ -H "Authorization: Bearer $NODUS_API_KEY"Rotate the signing secret
Section titled “Rotate the signing secret”Rotate when a secret may have leaked, or when you lost it:
$ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/rotations" \ -H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \ -H "Idempotency-Key: $(uuidgen)" -d '{"previousSecretTTL": "24h"}' | jq -r .status.signingSecretThe response holds the new secret, once. For previousSecretTTL (24 hours by default, at most 7 days) every
delivery carries a signature for the new secret and one for the old, so a receiver using either verifies it. Switch
your receiver to the new secret, and the old one stops signing when the time is up. status.secretRotation shows
when.
Check an endpoint’s health
Section titled “Check an endpoint’s health”$ nodus get webhookendpoint ci -o yamlstatus shows the latest delivery (lastDelivery), the deliveries of the last 24 hours that succeeded or failed
(deliveries24h), and, when the last three deliveries in a row failed, failingSince with the Healthy condition
set to False and reason DeliveriesFailing. Nodus also records a webhookendpoint.deliveries_failing event
when that happens, which you can send to another endpoint.