Hold every AI run to your analysts’ standard.
Your analysts write the bar for an earnings review or an IC memo. Stellifai tests every agent run against it, keeps the history of every rule, and finds the cheapest model that still clears it.
Book an eval sprint- Rule 1Revenue ties to the 10-Q within 0.5%
- Rule 2Guidance change in first paragraph
- Rule 3Every figure cites its source filing
- Rule 4Peers on the same fiscal calendar
Example data
Q3 earnings review · rubric v14
Rules:
1. Revenue ties to the 10-Q within 0.5%
2. Guidance change flagged in the first paragraph
3. Every figure cites its source filing
4. Peers on the same fiscal calendar
Agent log:
→ reading 10-Q
→ drafting earnings summary
→ checking rules 1–4
Models:
Model A$1.48/run
Model B$0.62/run
Model C$1.96/run
Example: run 0412 clears the bar on Model B at $0.62 per run.
The right rules. The right model.
A navigator needs a fix on every star before trusting a position. You need three things before trusting an agent’s work.
Write the bar
An analyst writes what a good earnings review looks like, rule by rule, in plain language. No engineer in the loop.
rule 2 · written by the coverage lead Flag any change to guidance in the first paragraph, with the prior figure.
Test every run
Each agent run is checked against every rule before it reaches a reviewer. You see which rule it missed and why.
run 0413 4 of 4 rules clears run 0414 4 of 4 rules clears run 0415 3 of 4 rules held: rule 4
Pick the model
When a new model ships, re-run the suite. Each workflow moves to the cheapest model that still clears the bar.
Model A $1.48/run clears Model B $0.62/run clears selected Model C $1.96/run clears
Example data
Every rule has a history.
Standards change. A PM tightens a tolerance, a new analyst adds a check. Every rule is versioned: who changed it, when, why, and whether last quarter’s runs would still clear it.
One team’s proven rubric becomes a starting point the next team can fork, adapt and re-test, so good standards spread without drifting.
TRIP LOGQ3 earnings review
Rule 1 tolerance tightened
- ties to the 10-Q within 1.0% + ties to the 10-Q within 0.5%
212 past runs re-tested · 9 no longer clear
Rule 4 added
+ peers on the same fiscal calendar212 past runs re-tested · 14 held on rule 4
Adapted for consumer coverage
~ 3 rules kept, 1 rewrittenbaseline set
Example data
Built for investment research.
| Destination | What it has to get right | Re-run |
|---|---|---|
| Earnings and coverage | Reviews that have to run the same way every quarter. | every quarter |
| Diligence | Where the numbers have to tie across every team's inputs. | every deal |
| Portfolio monitoring | One workflow, re-run against every asset you hold. | every asset |
Piloted with three institutional research and investing teams.
Start with one workflow.
In two weeks your team has a written bar, a test suite, and a cost per run for every model that clears it.
- Duration
- 2 weeks
- Scope
- 1 workflow
- You get
- written bar · test suite · cost per run by model