Datasets & training
The model can only learn from the examples you supply. Bring representative, varied traffic from the endpoint you intend to protect.
Choose an input format
Request-event JSONL contains one request object per line. The importer extracts configured fields, builds histories, and computes feature vectors. A line can look like this:
{
"timestamp_ms": 1000,
"endpoint_id": "22222222-2222-4222-8222-222222222222",
"method": "POST",
"path": "/orders",
"headers": { "X-Client-ID": "client-a" },
"query": {},
"client_ip": "192.0.2.1",
"body": { "amount": 42 },
"body_size": 13
}
Use the endpoint ID and template generated by your saved pipeline. The example ID above is not a shortcut for a real import. body_size must be the actual original body length. Requests before the longest-window warm-up completes contribute to history but do not become training vectors.
Download the example request events
Feature-vector CSV contains already-computed complete numbers. Its header names and order must exactly match the saved pipeline's feature list:
request_count,amount_mean,body_bytes
8,42.5,32
6,38.0,29
Upload the matching import manifest with its pipeline_version_id and schema_hash. PragmaChange can check a CSV's shape and finite float32 values, but cannot prove an outside aggregator used the right window semantics.
Download the example feature vectors
Offline recordings and custom generators
You can generate request JSONL or feature CSV yourself, without using a client library. The optional reference clients also write request JSONL with a finalized SDK recording manifest. Select both files in Datasets; the server verifies their checksum, counts, and saved pipeline binding. Those recording files have a stricter 4 MiB limit.
Behavior recordings are a different format and import into Behavior, where they become immutable snapshots for sequence training. They do not upload as endpoint forest data. See offline recordings.
Preview before accepting
Download the format template from the portal, upload your file, inspect the preview, and fix every row error. A dataset with any invalid row is not imported. The preview shows up to 100 row errors. Imports need at least 20 usable vectors. Source files are limited to 100,000 rows and a 15 MiB file picker limit within the 16 MiB HTTP request limit.
The downloadable examples above show the formats, but their sample IDs and hashes must be replaced with the ones from your saved pipeline. Never change a hash just to make incompatible vectors appear valid.
Imported source files and matrices live in the private artifact volume. The database holds metadata and file references; back them up together. The Nginx runtime does not automatically capture live production requests or delete endpoint imports. Application-side reference recorders collect only explicitly selected data; the behavior service has separate retention and purge rules.
What training does
The Python trainer fits an Isolation Forest to the first 80% of vectors and evaluates the last 20% as an unlabeled holdout. Training uses all ordered features, one worker, a fixed seed, no bootstrap, and bounded tree and sample settings. Rust validates the exported, data-only forest and checks score parity before publishing it. Models are not Python pickle files loaded in Nginx.
The holdout tells you how the model scores unused normal examples. Without labeled attacks and normal production outcomes, it cannot measure detection accuracy or a false-positive rate for your deployment. Replay realistic requests and tune thresholds deliberately.
Keep the training distribution honest
- Train each new pipeline on its own endpoint's normal behavior.
- Include normal variation in client activity, times, amounts, and payload sizes.
- Keep the dataset format tied to the saved pipeline version and schema hash.
- Rebuild the dataset and model after changing extractors, partitions, windows, or feature order.