# Webhooks

A long-running request can push its result to you instead of being [polled](/docs/async-jobs) for it. Pass a `webhook` object when you submit, and Layer `POST`s each of that request’s terminal events to your URL. The submit response’s [expected\_events](#expected%5Fevents) says which deliveries to wait for: usually the run’s own outcome, plus a separate scoring event where the workspace scores automatically.

Callbacks are opted into **per request**: there is no subscription to configure. The only workspace-level setting is the [signing secret](#verifying-a-delivery).

```bash
curl -X POST https://api.app.layer.ai/api/v2/workspaces/$WORKSPACE_ID/inferences \
  -H "Authorization: Bearer $LAYER_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
        "base_model_id": "BASE_MODEL_ID",
        "prompt": "a mossy stone golem",
        "webhook": {
          "url": "https://example.com/hooks/layer",
          "metadata": { "job": "hero-splash-42" }
        }
      }'
```

## Requesting a callback

| Field    | Required | Meaning                                                                                                           |
| -------- | -------- | ----------------------------------------------------------------------------------------------------------------- |
| url      | yes      | Public **HTTPS** endpoint that receives the POST. Must not redirect.                                              |
| events   | no       | Narrows what is delivered. Omit to receive every event the request can produce.                                   |
| metadata | no       | Opaque string key/values echoed back verbatim in every delivery. At most **10 keys** and **1024 bytes** in total. |

`http://`, private addresses and `localhost` are rejected at submit time with a [422](/docs/errors), so use a tunnel while developing.

These endpoints accept a `webhook`:

| Endpoint                               | Events it can produce                                                                                                                |
| -------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------ |
| POST .../inferences (**v2 only**)      | inference.completed, inference.failed, inference.cancelled: plus the scoring events below when the scope has automatic scoring rules |
| POST .../training-runs                 | training.succeeded, training.failed                                                                                                  |
| POST .../scores                        | scoring\_run.completed, scoring\_run.failed                                                                                          |
| POST .../workflows/{workflow\_id}/runs | workflow\_run.succeeded, workflow\_run.failed, workflow\_run.cancelled                                                               |

Caution

`POST /v1/.../inferences` is poll-only. Move to [/v2](/docs/migration) to be pushed generation results; the other three endpoints take a `webhook` on both versions.

### `expected_events`

Every response to those endpoints carries `expected_events`: the deliveries **this** request will produce. Wait for those and no others: a generation in a workspace with automatic scoring rules produces a second, separate event for its scores, and one without produces nothing further.

```json
{
  "inference_id": "…",
  "status": "in_progress",
  "expected_events": ["inference.completed", "inference.failed", "inference.cancelled"]
}
```

It is best-effort: the scoring rules are read at submit time, so someone disabling one mid-run can still leave a listed scoring event undelivered. Treat an event that is _not_ listed as a surprise to tolerate, not an error.

## The payload

Every delivery has the same envelope, and `data` is exactly what the matching `GET` endpoint returns for that job: no follow-up request is needed.

```json
{
  "id": "evt_7c9e6679…",
  "type": "inference.completed",
  "created_at": "2026-08-25T09:41:02Z",
  "data": { "inference_id": "…", "status": "complete", "outputs": [] },
  "metadata": { "job": "hero-splash-42" }
}
```

`data` only ever gains fields. Ignore ones you do not recognise rather than rejecting the body.

Terminal outcomes only: a job produces exactly one of its events, once. The per-event payload schemas are in the REST reference under **Webhooks**.

## Verifying a delivery

Deliveries carry [Standard Webhooks](https://www.standardwebhooks.com) headers, so any Standard Webhooks library verifies them with no code of your own:

| Header            | Meaning                                             |
| ----------------- | --------------------------------------------------- |
| webhook-id        | The event id, stable across retries.                |
| webhook-timestamp | Unix seconds when the attempt was signed.           |
| webhook-signature | One or more space-delimited v1,<base64> signatures. |

Verifying by hand: HMAC-SHA256 over `{webhook-id}.{webhook-timestamp}.{raw body}`, keyed by the base64-decoded secret (without its `whsec_` prefix), compared in constant time. Sign the **raw request body**: parsing and re-serializing produces different bytes and will not verify.

Reject deliveries whose `webhook-timestamp` is far from now (a few minutes’ tolerance), so a captured delivery cannot be replayed at you later.

### The signing secret

```bash
# Read the workspace's current secret (minted on first read)
curl https://api.app.layer.ai/api/v1/workspaces/$WORKSPACE_ID/webhook-secret \
  -H "Authorization: Bearer $LAYER_TOKEN"
# → { "secret": "whsec_…" }


# Rotate it
curl -X POST https://api.app.layer.ai/api/v1/workspaces/$WORKSPACE_ID/webhook-secret/rotate \
  -H "Authorization: Bearer $LAYER_TOKEN"
# → { "secret": "whsec_…", "previous_secret_expires_at": "…" }
```

Both endpoints need workspace **admin** rights. The secret is a credential: anyone holding it can forge deliveries into your receivers.

A rotation puts the outgoing secret into a grace window: until `previous_secret_expires_at`, both secrets sign every delivery, so you can update a receiver without dropping anything. That is why more than one signature can appear in the header: a delivery is authentic if **any** signature matches.

## Responding

Answer with any 2xx within **10 seconds**. The body is ignored. Acknowledge first and do your work afterwards; a receiver that finishes its processing before answering will time out.

| Your response               | What happens                                                          |
| --------------------------- | --------------------------------------------------------------------- |
| Any 2xx                     | Delivered.                                                            |
| Any non-2xx, or no response | Retried.                                                              |
| 429                         | Retried, honouring Retry-After (in seconds) where it asks for longer. |
| 410 Gone                    | Delivery stops permanently. This is how you retire an endpoint.       |
| 3xx                         | Counts as a failure: redirects are never followed.                    |

## Retries and duplicates

Five attempts over roughly an hour: immediately, then after \~10s, \~1m, \~10m and \~45m. If they are all exhausted nothing is lost: the job’s `GET` endpoint returns the same body for as long as the job exists, which is what makes a short retry window safe.

Delivery is **at-least-once and unordered**. Deduplicate on `webhook-id`, which does not change between attempts of the same event, and use `created_at` to order events for one job. In particular `inference.completed` and its `scoring_run.completed` can arrive in either order.

Every attempt is recorded on the job itself: `GET` the run and read its `webhook` block for `status`, `attempts`, `last_attempt_at`, `last_response_status` and `last_error`. That block is the only place a failed delivery is visible.

## Before you go live

`POST .../webhook-tests` sends a signed sample to any URL and tells you synchronously what came back: reachability, latency, status code. It costs nothing and creates no run, so it is the cheapest way to check a tunnel or a signature check before spending Creative Units.

```bash
curl -X POST https://api.app.layer.ai/api/v1/workspaces/$WORKSPACE_ID/webhook-tests \
  -H "Authorization: Bearer $LAYER_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{ "webhook": { "url": "https://example.com/hooks/layer" } }'
# → { "delivered": true, "response_status": 200, "latency_ms": 143 }
```

The sample carries `type` `inference.completed` with an empty `data` object and is signed exactly as a real delivery is.

Note

Webhooks tell you a run **settled**, not that it succeeded. A failure arrives through the same callback as a success, including one that submission could not predict, such as `error_code: INSUFFICIENT_BALANCE`. Inspect `data.status` and `data.error_code` as you would when polling. See [Errors](/docs/errors).
