AI configuration
Ship a prompt change like any other rollout.
Manage prompts, models, and parameters as versioned config on a flag. Ramp a new model to 10%, measure quality and cost against the current one, and roll back if it regresses — with the LLM spans sitting right next to the readout.
ProIncluded from the Pro tier.
assistant-modelA/B · gpt vs. nextA prompt change, measured like a feature.
What you get
Why Foredeck for ai configs.
Versioned prompt & model config
Store prompts, model ids, and parameters as flag config so a change is a ramp, not a redeploy.
Measured against production
Treat the AI change as an experiment: quality, latency, and token/cost metrics compared to the current config.
LLM traces alongside
GenAI spans (model, tokens, tool steps) land next to the experiment that ramped them, so you can see cost and latency per call.
Guardrailed
Put a cost or quality guardrail on the ramp and let auto-rollback stop a bad prompt.
How it works
Three steps, one pipeline.
Put the prompt/model on a flag as config variations.
Send a slice of traffic to the new config and emit exposures.
Read quality/cost/latency lift and promote or roll back.
Better together
Works with the rest of the platform.
Experiments
FreeEvery rollout, automatically measured — experiments included free.
Explore Experiments →Tracing
FreeFollow one call across every service — spans stamped with the variation.
Explore Tracing →Guardrails & Auto-Rollback
ProCatch a regressing release before it spreads — automatically.
Explore Guardrails & Auto-Rollback →Metrics
FreeDefine a metric once, use it everywhere.
Explore Metrics →One object, one pipeline. The same rollout you ship is automatically measured, guarded, and — when it breaks — traceable, because every signal is joined to the flag and variation the unit was in.
Get in