DriftGate turns production traces into calibrated regression tests and detects statistically meaningful behavioral drift.
When developers tweak system prompts (e.g. asking an agent to "be more concise and empathetic"), standard evaluations miss subtle regressions. Output schemas mutate silently, required keys vanish, and uncalibrated LLM judges rubber-stamp broken responses.
A self-calibrating pipeline that continuously converts production telemetry into rigorous, deterministic regression test suites.
Raw multi-turn logs and agent tool calls are normalized into strict Pydantic schemas and embedded into 1536-dimensional vector branches using Neon Serverless Postgres.
HDBSCAN automatically discovers true operational behavior boundaries across embedding space without requiring manual annotations or hardcoded categories.
A 4-node LangGraph agent analyzes cluster histories to generate typed assertion verifiers (regex patterns, JSON schemas, and semantic invariants).
Before trusting a test, DriftGate tests the test. Candidate verifiers are validated against golden ground-truth slices. Rubber-stamp tests (FPR > 0.10) are hard-purged.
40 parallel test cases per cluster execute across Baseline (Prompt v1) and Candidate (Prompt v2) on Featherless AI with SHA-256 caching.
Computes Wilson 95% Score Confidence Intervals and Benjamini-Hochberg FDR control across all clusters, mathematically proving true behavioral regressions.
Inspect synthesized verifier code, AST constraints, and frequentist scorecards in real time.
Zero code rewrites. Connect DriftGate with any orchestration framework, proxy layer, or local inference gateway.
DriftGate ingests raw multi-turn customer logs and tool trajectories, canonicalizing unstructured traces into typed Pydantic models with 1536-d Neon vector branches.
DriftGate turns production traces into calibrated regression tests and statistically identifies meaningful behavioral drift before it reaches users.