[ Methodology ]
How we'll know if Jonbar is right.
A rehearsal is only worth something if it's been checked against reality. This is how we check, and what we'll publish.
Status: the blind backtest is being built. Until it's published, every run says "not yet backtested".
[ 01 ] Test case
What a test case is
A real, past commercial decision with a public outcome: a price change, a campaign mechanic, a launch price. It needs at least two options (what was done and the alternative) and a reported result on a named metric (units, revenue, margin).
[ 02 ] Blind order
Blind, in this order
- 01Casebefore | outcome
- 02Sealoutcome → hash
- 03Runfixed engine version
- 04Lockversion + seeds committed
- 05Scoreabstain = miss
- 01
Split and seal.
The case is split in two: what was known before the decision, and the outcome. The outcome is sealed; only its fingerprint (a hash) is stored.
- 02
Run on the "before" part.
The engine runs on the "before" part only, at a fixed version.
- 03
Lock the prediction.
The prediction is committed with its version and seeds before anyone opens the outcome.
- 04
Open and score.
We unseal the outcome and score direction and size of the miss. Any change to the engine after that is a new iteration, tested on new sealed cases.
- 05
"I can't tell" counts as wrong.
You can't score well by skipping hard cases.
[ 03 ] Two numbers
Two numbers, always together
- Direction: did we pick the option that actually did better?
- Size of the miss: how far our predicted effect was from the real one, in percentage points.
Neither is shown alone. Direction on its own is a weak bar: even simple rules often get it right.
[ 04 ] Baselines
Against simple rules
- The textbook rule: price up → units down; discount up → sales up.
- A fixed average elasticity.
- A plain AI language model asked "which option won?", with no simulation (planned; measured once enabled).
We report Jonbar's score next to theirs on the same cases. We only say "better" if the gap is statistically significant.
[ 05 ] Sample size
Honest about sample size
The first sealed test set will be small, so the first number will be a rough signal, not proof. We say so on the result, with its range.
[ 06 ] Your campaigns
Your own campaigns, blind
In a pilot we pick one of your past campaigns, see only data from before it started, lock the prediction, and then compare it with what really happened. This is also how accuracy on Turkish marketplace data gets measured: it hasn't been yet.
[ 07 ] Ranges
What the ranges mean
Each run shows a range for its result. That range is the spread under the model's own assumptions, not a probability that reality lands inside it. We'll tell you how often real results fell inside these ranges once the backtest is published.
[ 08 ] Publishing
What gets published
The result, good or bad: direction score, median miss, how often real results fell inside the ranges, the simple-rule scores, and every high-confidence miss listed by name of case.
[ 09 ] The market model
The whole market, layer by layer.
A decision meets shoppers, a shelf of competitors, a platform's fees and its ranking. We're building them as separate layers, because each one is calibrated with different data, and we open a layer only when there is data that could prove it wrong.
- CalibrationBuilding
Every forecast will be locked at decision time and compared with the real result, so the model can be corrected.
- Ranking algorithmRoadmap
How the platform's ranking responds to price, campaigns and ratings. Not modelled today; it needs ranking history to calibrate.
- Platform economicsBuilding
Commission, shipping or courier share, the discount the platform funds and product cost will come off each order. Output: contribution margin.
- ShelfBuilding
The competing list at the moment you decide, taken as one snapshot: prices, ratings, campaigns. How competitors react will stay your assumption.
- Shoppers (demand)In pilot
Demand fitted to your own sales: price response, promo lift, pull-forward, the dip after a campaign, cannibalization.