Browse documentation ↓
04 / DEFINE THE SIGNAL

Pipelines & features

A pipeline turns request details and recent activity into an ordered list of numbers for one endpoint.

Build from four pieces

  1. Endpoint: one API host, HTTP method, and path template.
  2. Extractors: names for request fields you need, such as a header, query value, path parameter, JSON Pointer, or client IP.
  3. Partitions and windows: who gets separate history, and how far back each measurement looks.
  4. Features: ordered numeric outputs. The order is part of the trained model contract.

Only extract fields used by a feature or partition key. Required fields must be present when replaying or serving requests. Header names are case-insensitive; other extracted strings are case-sensitive. Missing fields, JSON null, objects or arrays where scalars are expected, duplicate selected headers or query parameters, and invalid numeric values cause extraction errors.

Feature operations

OperationMeaning
count, rateNumber of requests, or count divided by the full window length in seconds.
sum, mean, min, max, stddevNumeric values observed in a window. Standard deviation uses the population formula.
distinct_countExact count of distinct typed scalar values in a window.
interarrival_mean, interarrival_stddevGaps between consecutive request arrivals inside the window, in milliseconds.
current_valueNumeric value from this request, without a window.
body_sizeCurrent request body size in bytes, without a window.

Windows use (now - duration, now]: an event exactly on the left boundary is excluded. In production the clock is monotonic processing time. Offline request replay uses supplied timestamps in sorted order and keeps source order for ties.

Example: orders by client

This is a shape example from the product repository. IDs and host are illustrative; save a real pipeline in the portal before importing data.

{
  "api_host": "api.example.test",
  "method": "POST",
  "path_template": "/orders",
  "extractors": [
    { "name": "client", "source": { "kind": "header", "name": "X-Client-ID" } },
    {
      "name": "amount",
      "source": { "kind": "json_pointer", "pointer": "/amount" }
    }
  ],
  "partitions": [{ "name": "by_client", "components": ["client"] }],
  "windows": [{ "name": "last_second", "duration_ms": 1000 }],
  "features": [
    {
      "name": "request_count",
      "operation": "count",
      "source": null,
      "partition": "by_client",
      "window": "last_second"
    },
    {
      "name": "amount_mean",
      "operation": "mean",
      "source": "amount",
      "partition": "by_client",
      "window": "last_second"
    },
    {
      "name": "body_bytes",
      "operation": "body_size",
      "source": null,
      "partition": null,
      "window": null
    }
  ]
}

The full checked example, including IDs and limits, is examples/pipeline.json in the product repository. The portal fills in saved IDs and hashes for its templates.

Download the example pipeline JSON

Versions and compatibility

Saving creates an immutable pipeline version. A model binds the version ID, schema hash, and ordered feature names. A changed pipeline needs a new dataset and model and gets separate state. A model-only update using the same pipeline ID can retain compatible window history.

New pipelines have one endpoint. Previously saved groups retain their original history rules and release compatibility, but you cannot create a new group through the v1 panel or API.

Assisted setup can propose these definitions for review. Reference recorders use the saved pipeline binding; they do not generate pipelines or precompute the native window features.