Experiments & feature flags

Shipping on a hunch is a decision. It just isn’t a measured one.

Flags and A/B tests in the same system as your product metrics, so the variant that wins is judged on activation and retention — not on clicks counted by a tool that has never seen your funnel.

  • One call to flag a change
  • Rollback in seconds
  • Replaces LaunchDarkly and Optimizely
Sound familiar?

Every team runs experiments. Fewer can say what they proved.

“It tested well.”

Well against what? The flag vendor measured clicks, and clicks are the one metric a worse experience can win.

“We will measure it after we ship.”

Which means the tracking gets written in a hurry, by someone else, after the variant is already live to everyone.

“Can we turn it off right now?”

The rollback is a deploy, the deploy is a queue, and the incident is measured in the minutes between.

Experiment · pricing page50 / 50 split⚠ Not yet significant — do not call itB − A+0.0 pts
AControl00.0%
Flag → 0% traffic · no deploy
BVariant B00.0%
Flag → 0% traffic · no deploy
Significance reached · p < 0.01

The control led at halfway.
It still lost.

Calling it early would have shipped the wrong page. When it did resolve, the losing variant went to zero traffic from the flag — no deploy, no rollback ticket.

What you get

A test is only worth running if the result can change your mind.

Measured on what matters

The result and the rest of your analytics, in one place.

Exposure and conversion land in the same store as your goals and funnels, so the variant that won on clicks can be read against activation and retention without exporting anything or reconciling two vendors.

  • Read next to the goals and funnels you already track
  • Significance reported plainly, without a stats degree
  • Corrected for the number of variants, so a third arm cannot fake a winner
onboarding_v3Significant

Measured against goal · Activation

Control25%
Variant B32%
+7 pts activation · retention heldRoll out
Rollback

Off in seconds, not at the next deploy.

Turning a flag off invalidates the served bundle immediately, so the next page view anywhere in the world reads the new state within seconds. The kill switch is a toggle rather than a release.

  • Percentage rollouts, held back per plan, per segment or per account
  • Version-key polling — no long-poll connection to keep alive
  • No-code changes for copy and layout variants
Variant performanceLast 28 days
Control100%
Variant A64%

Biggest drop — 36% lost here · watch the friction sessions

Variant B41%
Rolled out7%
The agent

It proposes the test, and writes the tracking.

Ask what to try next and the agent reads where the funnel leaks, ranks the experiments worth running against it, drafts the variant and wires the measurement. A person still approves everything that ships.

  • Tests proposed from where the funnel actually leaks
  • Tracking written for you, before the variant goes live
  • Pull requests open for review — nothing merges itself
Ask FlowgridLive

Did the new onboarding actually beat the old one?

Answer
  • Control96%
  • Variant A70%
  • Variant B46%
  • Rolled out18%
Working shownFollow-ups keep contextNo SQL
Guardrails

Moving fast is only fast if you can undo it.

onboarding_v3Live · 50%
Serve variant B
version keyv44
clients on old experience0%
DecisionServed

Illustrative — the bundle is invalidated on the flip, so undoing it does not wait for a deploy.

Nothing ships itself

The agent can open a pull request. Merging it is a person, every time, with a preview of the change and the approval recorded.

Flags fail safe

If the flag service cannot be reached, your code falls back to its default. A bad network is never a surprise rollout.

One definition of winning

Experiments read the same goals as your dashboards, so "it won" means the same thing in the test as it does in the board review.

What changes
Seconds
From deciding to roll back to your users seeing the old experience — served fresh on the next page view, not shipped by deploy.
One call
To read the variant and render your branch — and one more to record the conversion you are testing against.
One verdict
Significance is computed server-side on a schedule, so the winner badge and the Ship button always read the same numbers.
The maths

Flags and metrics, or two vendors and a reconciliation.

Flowgrid
A flags vendor plus your analytics
Feature flags and rollouts
Included
Your flag vendor
A/B tests on the same events
Same system, same goals
A second tool to reconcile
Measured on retention, not clicks
Your goals, recomputed daily
Whatever the tool counts
Knowing when to ship
Ship stays locked until the result clears
A judgement call in a meeting
Writing the tracking
The agent drafts it, you approve
A ticket in somebody else's queue
Rollback speed
The next page view
The next deploy

Most teams arrive paying a flags vendor and an analytics tool that cannot see each other, which is why the results argument never quite closes.

Before you ask

What engineering will ask in review.

One subscription in place of the stack you're running today

MixpanelAmplitudeHotjarFullStoryLaunchDarklyOptimizelyChartMogulKlipfolioTriple WhaleGainsightVWOPlausible

Flags should not cost more than the feature.

Experiments and feature flags are included on every paid plan, next to the analytics that decide whether the variant actually won.

  • Free to start, no credit card
  • Unlimited flags on every paid plan
  • Run it alongside your current vendor while you migrate
See pricing