Skip to content

A WebhookEndpoint sends an HTTPS request to your server each time an event you care about happens: a Job succeeds or fails, a balance runs low, a Budget crosses a threshold. You can have as many endpoints as you need, each with its own URL, event types and projects.

Every delivery is signed with the Standard Webhooks scheme, so you can check that it came from Nodus and was not changed in transit.

endpoint.yaml
apiVersion: nodus.dev/v1
kind: WebhookEndpoint
metadata:
name: example-receiver
spec:
url: https://hooks.example.com/nodus
description: Job results for the example receiver
eventTypes:
- job.succeeded
- job.failed
- billing.low_balance
Terminal window
$ nodus apply -f endpoint.yaml -o json | jq -r .status.signingSecret
whsec_...

The signing secret is in the response to the create, and nowhere else. Copy it into your receiver’s configuration now. If you lose it, rotate it to get a new one.

Nodus sends a webhookendpoint.ping delivery right after the create, so you can see your receiver work before any real event.

Field Meaning
url Where deliveries go. https only, with a public address. No redirects are followed.
eventTypes Event types to send: an exact type (job.succeeded), a prefix (job.*, billing.*) or *. Defaults to *.
projects Send project events only for these projects. Empty means every project. Org-level events always match.
enabled false pauses deliveries. Pending deliveries to a paused endpoint are dropped.
description A note for you.

Every field can change after the create:

Terminal window
$ nodus patch webhookendpoint ci --type merge -p '{"spec":{"eventTypes":["job.*"]}}'

A delivery is a POST with a JSON body:

{
"type": "job.succeeded",
"timestamp": "2026-10-01T12:00:00Z",
"data": {
"event": { "reason": "Succeeded", "message": "…", "count": 1 },
"object": {
"apiVersion": "nodus.dev/v1",
"kind": "Job",
"metadata": { "name": "train", "namespace": "default", "uid": "job_…" },
"status": { "phase": "Succeeded" }
}
}
}

and three headers:

Header Meaning
webhook-id Names this delivery. It stays the same across retries, so use it to ignore duplicates.
webhook-timestamp When the attempt was signed, in Unix seconds.
webhook-signature One or more signatures, separated by spaces, each like v1,<base64>.

The event type is <kind>.<reason> in lowercase with underscores, such as job.succeeded or agentrun.needs_resolution. Billing events use the billing. prefix.

Type When
job.succeeded, job.failed, job.cancelled A Job reached a final phase. The same reasons exist for other run kinds, such as agentrun.failed.
job.preempted, job.node_lost, job.restored A Job lost its machine, and was restored from its last checkpoint.
job.checkpointed, job.checkpoint_failed A checkpoint was saved, or could not be.
job.max_cost_reached, job.funding_lost, job.funding_restored A Job stopped or resumed because of its cost cap or your balance. The same reasons exist for Sandboxes and other kinds with a cost.
job.output_sink_failed A Job’s output could not be loaded into its destination.
job.gang_restarted, job.network_degraded A multi-node Job restarted or lost network quality.
sandbox.started, sandbox.stopped, sandbox.stuck, sandbox.setup_failed A Sandbox changed state.
agentrun.needs_resolution An AgentRun is waiting for you to resolve a step.
billing.low_balance Your available balance fell below the warning level.
billing.auto_recharge_succeeded, billing.auto_recharge_failed An auto-recharge added credits, or its charge failed.
billing.arrears_posted A charge was larger than your balance, so new work waits until it is paid.
billing.dispute_opened, billing.dispute_closed A payment was disputed, which pauses new work, or the dispute closed.
billing.storage_quota_exceeded, billing.egress_quota_exceeded A quota was reached.
budget.threshold, budget.exceeded A Budget crossed a threshold, or refused a new run.
topup.succeeded, topup.failed, topup.refunded A top-up completed, its payment failed, or it was refunded.
pool.node_offline, pool.forecast_shortfall, pool.action_proposed Events of your own capacity pools.
webhookendpoint.ping Sent once when an endpoint is created.
webhookendpoint.deliveries_failing Your receiver has failed every recent delivery.

nodus get events shows the events of your org; each webhook type is the kind and reason of one of them.

Check every delivery before you trust it. The signature covers the exact bytes of the body, so verify the raw body before you parse it. Use a Standard Webhooks library for your language; it also rejects timestamps more than five minutes old, which stops a captured delivery from being replayed later.

main.go
package main
import (
"encoding/json"
"fmt"
"io"
"log"
"net/http"
"os"
"sync"
"time"
standardwebhooks "github.com/standard-webhooks/standard-webhooks/libraries/go"
)
// maxBody bounds a delivery; Nodus sends at most 64 KiB.
const maxBody = 64 << 10
// event is the body of a delivery: the event's type and the Event and object it is about.
type event struct {
Type string `json:"type"`
Timestamp time.Time `json:"timestamp"`
Data struct {
Event json.RawMessage `json:"event"`
Object struct {
Kind string `json:"kind"`
Metadata struct {
Name string `json:"name"`
Namespace string `json:"namespace"`
} `json:"metadata"`
} `json:"object"`
} `json:"data"`
}
// receiver verifies and handles deliveries.
type receiver struct {
hook *standardwebhooks.Webhook
out io.Writer
mu sync.Mutex
seen map[string]bool
}
func newReceiver(secret string, out io.Writer) (*receiver, error) {
hook, err := standardwebhooks.NewWebhook(secret)
if err != nil {
return nil, fmt.Errorf("NODUS_WEBHOOK_SECRET: %w", err)
}
return &receiver{hook: hook, out: out, seen: map[string]bool{}}, nil
}
// ServeHTTP answers 2xx once the delivery is verified and recorded, which is all Nodus needs to stop retrying. A
// delivery that fails verification gets 400 and is never parsed: the signature covers the exact bytes received, so
// verify before you decode.
func (r *receiver) ServeHTTP(w http.ResponseWriter, req *http.Request) {
body, err := io.ReadAll(http.MaxBytesReader(w, req.Body, maxBody))
if err != nil {
http.Error(w, "body too large", http.StatusRequestEntityTooLarge)
return
}
// Verify checks the signature against every secret in the header and rejects timestamps more than five minutes
// from now, so a captured delivery cannot be replayed later.
if err := r.hook.Verify(body, req.Header); err != nil {
http.Error(w, "invalid signature", http.StatusBadRequest)
return
}
// A delivery can arrive more than once (a retry after a slow answer): webhook-id names the delivery, so handle
// each id once and still answer 200.
if r.firstTime(req.Header.Get(standardwebhooks.HeaderWebhookID)) {
var e event
if err := json.Unmarshal(body, &e); err != nil {
http.Error(w, "malformed event", http.StatusBadRequest)
return
}
r.handle(e)
}
w.WriteHeader(http.StatusOK)
}
func (r *receiver) firstTime(id string) bool {
r.mu.Lock()
defer r.mu.Unlock()
if r.seen[id] {
return false
}
r.seen[id] = true
return true
}
// handle does the work of an event. Do slow work after answering (queue it); Nodus waits 15 seconds for a response.
func (r *receiver) handle(e event) {
obj := e.Data.Object
fmt.Fprintf(r.out, "%s %s/%s\n", e.Type, obj.Kind, obj.Metadata.Name)
}
func main() {
rcv, err := newReceiver(os.Getenv("NODUS_WEBHOOK_SECRET"), os.Stdout)
if err != nil {
log.Fatal(err)
}
srv := &http.Server{Addr: ":8080", Handler: rcv, ReadHeaderTimeout: 5 * time.Second}
log.Fatal(srv.ListenAndServe())
}

This receiver runs as is: NODUS_WEBHOOK_SECRET=whsec_... go run ./examples/webhooks/receiver. The same check in Python:

from standardwebhooks.webhooks import Webhook
wh = Webhook(secret) # the whsec_ secret
event = wh.verify(request_body, request_headers) # raises if the signature or timestamp is wrong

If you cannot use a library, compute base64(HMAC-SHA256(key, id + "." + timestamp + "." + body)), where key is the secret after the whsec_ prefix, base64-decoded, and compare it in constant time with each v1, signature in the header.

A 2xx response is success. Anything else, a timeout or a connection error is retried with exponential backoff, starting at 30 seconds and doubling up to 8 hours between attempts, for 3 days after the event. After that the delivery is marked failed. A URL that Nodus is not allowed to connect to, such as a private address, fails at once.

Deliveries can arrive more than once and out of order. Use webhook-id to ignore a duplicate, and the timestamp in the body or the object’s status to see which state is newer.

Nodus keeps the log of an endpoint’s deliveries for 7 days:

Terminal window
$ curl -s "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries?limit=20" \
-H "Authorization: Bearer $NODUS_API_KEY"

Each entry shows the event type, the number of attempts, the receiver’s last response code, how long it took, and a phase: Pending while retries remain, then Succeeded or Failed. A delivery to an endpoint that was deleted or disabled is Abandoned. The response includes continue when there is more; pass it back as after.

To send a delivery again, replay it. The replay has the same body and a new webhook-id:

Terminal window
$ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/deliveries/$ID/replays" \
-H "Authorization: Bearer $NODUS_API_KEY"

Rotate when a secret may have leaked, or when you lost it:

Terminal window
$ curl -s -X POST "$NODUS_API_URL/apis/nodus.dev/v1/webhookendpoints/ci/rotations" \
-H "Authorization: Bearer $NODUS_API_KEY" -H "Content-Type: application/json" \
-H "Idempotency-Key: $(uuidgen)" -d '{"previousSecretTTL": "24h"}' | jq -r .status.signingSecret

The response holds the new secret, once. For previousSecretTTL (24 hours by default, at most 7 days) every delivery carries a signature for the new secret and one for the old, so a receiver using either verifies it. Switch your receiver to the new secret, and the old one stops signing when the time is up. status.secretRotation shows when.

Terminal window
$ nodus get webhookendpoint ci -o yaml

status shows the latest delivery (lastDelivery), the deliveries of the last 24 hours that succeeded or failed (deliveries24h), and, when the last three deliveries in a row failed, failingSince with the Healthy condition set to False and reason DeliveriesFailing. Nodus also records a webhookendpoint.deliveries_failing event when that happens, which you can send to another endpoint.