// Portfolio

Forty tests running. Which three are lying?

Your testing platform shows you one activity at a time, and it tells you the same thing about each one: a conversion rate, a lift, a confidence number. It will not tell you that the traffic split on test 12 is broken, that the winner on test 5 stops being a winner once you count how many comparisons you ran, or that test 19 closes on Friday and was never going to reach an answer. Drop in one export covering everything you have live. Get every activity back, ranked worst first.

// Portfolio · ranked by severity
demo portfolio · 8 activities
Status
Activity
Raw lift
Honest lift
Split
Closes
Flags

Click any row for the full working. Every number is computed in your browser. Nothing is uploaded, and nothing is stored.

// Run it on yours

One file. Every activity.

Export your activities with a row per experience. It does not matter which platform: if the file has an activity name, an experience name, visitors and conversions, it will read them. Revenue and end dates unlock the revenue conflict and pace checks.

Drop a multi-activity CSV here
Adobe Target, Adobe Analytics, Optimizely, VWO, AB Tasty, Dynamic Yield, GA4, or a spreadsheet you built by hand. Columns are detected automatically and you can correct them.
Activity,Experience,Visitors,Orders,Revenue,End date PDP sticky add-to-cart,Control,48201,2410,289200,2026-08-04 PDP sticky add-to-cart,Sticky bar,45882,2477,300500,2026-08-04 Checkout progress bar,Control,61200,3060,367200,2026-07-28
// The argument

What your platform will not tell you

Every one of these is computable from the export you already have. None of them appear in the report your platform hands you, and none of them appear in the deck your agency builds from it.

CheckAdobe TargetOptimizelyAB Integrity
Sample ratio mismatch Did the traffic actually split the way you configured it? If not, nothing else in the test means anything. nopartialyes
Multiple-comparison correction Four experiences against five metrics is twenty chances to find a winner by luck. Both platforms report every cell as if it were the only one you looked at. nopartialyes
Achieved power The smallest lift this sample could ever have detected. Without it, "no significant difference" is indistinguishable from "we did not measure". nonoyes
Winner's-curse correction A barely significant winner is nearly always smaller than it looks. This is why shipped tests underdeliver against the deck six months later. nonoyes
Conversion against revenue More orders at a smaller basket is a loss wearing a win's clothes. Reported side by side, never reconciled. nonoyes
Pace to close At today's traffic, will this test reach an answer before its end date? Nobody computes this, so tests close on the calendar instead of on the data. nopartialyes
One ranked view across every activity Not a list of tests. A list of the tests that need you today, worst first. nonoyes

On the comparison. Adobe Target reports confidence from a two-tailed Welch's t-test as 1 minus the p-value, and its documented CSV download contains raw data only, without lift or confidence. Neither its reporting documentation nor Optimizely's describes a sample ratio check on the results view or a correction across experiences and metrics. Optimizely's sequential testing does control error under repeated looks and its export reports samples remaining to significance, which is why it scores partial on three rows rather than no. Marked as of July 2026 against published documentation. If a platform ships one of these, this table gets corrected.

On the method. Significance is a two-proportion z-test on the pooled standard error. The multiple-comparison bar is Šidák at a family-wise 5%, counting every variant against every metric in the file, which assumes the comparisons are independent and is therefore a touch conservative when your metrics correlate. Sample ratio mismatch is a chi-square goodness-of-fit test against the configured split, flagged hard below p = 0.01. The honest lift is a posterior mean under a skeptical normal prior of 15% relative on the true effect. Power is the smallest relative effect detectable at 80% with the sample collected. Every one of these is the same math as the free tools on this site, run across your whole portfolio at once instead of one test at a time.

// Access

The tools stay free. This one does not.

Every calculator on this site stays exactly where it is: free, no account, nothing uploaded. What costs money is not having to run them one test at a time. Upload once a week, get the whole portfolio triaged, and know which three tests to open before your Monday meeting.

Always free
€0
The seven tools, plus Instant Analysis on a single test. Everything runs in your browser.
  • One test at a time, by hand
  • Every calculation, in full
  • No account, nothing stored
See the tools
Portfolio · in development
€49 / month
Your whole portfolio, triaged from one file. Priced to go on a company card without a procurement conversation.
  • Unlimited activities per upload
  • Every check on this page, ranked worst first
  • Weekly change log: what moved, what broke
  • Shareable read-only verdict for the deck
  • Direct platform connections when they land
Ask for early access

The demo above is the real engine on invented data. Bring your own export and it will run on that instead, in your browser, before you pay anyone anything.