How to test AI quoting software, including ours
Software vendors do not publish their own exam. This is ours: the blind test we run before any deal, the pass bar written down in advance, how we score, and what we have got wrong so far.
Why a blind test instead of a demo
Every demo in this category works. That is what a demo is for. The vendor chose the document, knows where the traps are, and has driven the same path a hundred times. You learn how the software looks, not whether it works on your desk.
A blind test is different in one specific way: the vendor cannot see the answer. You take jobs you have already quoted and won or lost, hold back the filed quotes, and let the system price the packages cold. Then you compare against numbers you already know were right, because you lived them.
The result is a number that means something, and a conversation that is about evidence rather than adjectives. It also protects you from the most common failure in this market, which is not a vendor lying. It is a vendor measuring something easier than the thing you care about, and quoting that number in good faith.
The protocol, step by step
- Pick one product family. Not your whole portfolio. The family you quote most often, where the documents look alike enough that a pattern exists. A narrow deep slice tells you whether the thing works. A wide shallow one tells you nothing, twice as slowly.
- Gather roughly fifty package and quote pairs. The customer's bid package as it arrived, and your own answer to it: the quote you sent, and the costing sheet behind it if you have it. A package without your answer is half a case study.
- Seal about ten of them as the blind set. Chosen by project, not by convenience, and excluded everywhere the name appears. If a copy of a blind project sits in another folder, it is no longer blind. We do this exclusion in writing and you hold the list.
- We mine the rest. Your pricing logic, your rate history, your standards, your conventions. This is where the engine learns how your shop actually prices rather than how a textbook says it should.
- Run the blind set cold. The real engine, the one that would serve you in production, not a research setup. Before spending anything we verify that the blind folders contain no answer files, checked by name and by number, because a total like 346,668 sitting inside an instrument is a leak even with no name attached to it.
- Score together, live. Line by line, ours against yours. You see every miss, not a summary. This happens in one session so that nothing is quietly adjusted between the run and the report.
- Decide against the bar you wrote down. Pass or fail on the threshold agreed in step one. If it fails, you have spent weeks and learned something true, which is a considerably better outcome than a year of adoption you have to unwind.
How we score, and why it is two numbers
Accuracy is the most abused word in this market, because a single percentage can be made to mean almost anything. We report two, separately, and you should demand the same from anyone.
Reading accuracy
Of the requirements actually in the package, how many did the system extract, miss, invent, or correctly flag as ambiguous? Checked line by line against the documents, with every extracted value carrying the page it came from so you can verify it in a glance rather than a re-read.
This is scored first and separately, because reading is where the errors live. Every failure class we have logged while building this engine has been a reading, scope, classification or policy error. Not one has been arithmetic.
Pricing accuracy
On the blind jobs, how far is the total from the number you filed, and for the right reasons? Reported raw and again era-normalised, with the rate difference on its own line. Old quotes are priced in old money; aluminium and steel moved a long way across 2024 to 2026, and a system compared against three-year-old numbers without adjusting for that is being marked against a moving target.
We also bucket the gap so the buckets sum exactly to the total difference. A gap you cannot break down is a gap nobody understands, including us.
What we deliberately do not count
- Component accuracy. If a system is handed each block's own filed parts and then prices them, you have measured the calculator, not the product. The real number is end to end, package in, quote out. Any vendor quoting component accuracy as product accuracy will be caught by your engineer with one question, so it is better for everyone to be explicit about it up front.
- Correct totals from wrong steps. Two errors in opposite directions produce a total that looks excellent and a quote that is nonsense. We score derivations, not just the bottom line, which occasionally means reporting a worse-looking number than we could have.
- Anything scored after the fact. If the test set is visible before the run, the result is a measure of tuning rather than capability. This is why scoring is live and the blind list is yours.
The scorecard so far
Everything below runs on public documents, so you can download the same file and check us. Reading results are what we have. The pricing column is nearly empty and we explain why underneath rather than quietly leaving it out.
| Public package | Pages | What the engine returned | Time |
|---|---|---|---|
| Cape Vincent, New York. Waterfront improvements, gangways and floating docks | 179 | 214 statements extracted, each cited to its source page. 7 places where the drawings and the specification disagree, including two decking materials for the same structure and two deflection limits on the same page | 44 seconds |
| Hailey, Idaho. Water reclamation facility headworks | 668 | Found the 14 sections relevant to the equipment being quoted without being told where to look, and pulled the requirements from them with page references | 102 seconds |
| Grosse Ile, Michigan. Bar screen replacement | 2 | Named equipment and the demolition scope hiding in the same two pages. Honestly, a document this short does not need software | seconds |
| Odessa, Missouri. Mechanically cleaned bar screen | 4 | The sentence on page two that ends the exercise: substitute materials or work shall not be permitted | seconds |
| Miami-Dade, Florida. Access hatch standard specification | 12 | Requirements extracted from a standard specification section, the kind that gets incorporated by reference and rarely read | seconds |
The two short packages are in the table for honesty, not as evidence. A four page invitation does not need software and we are not going to pretend otherwise. They earn their place because one of them contains the single most decisive sentence in the set.
Why the pricing column is nearly empty
To score pricing you need a known right answer, and public bid packages almost never come with one. The specification tells you what was wanted. It does not tell you what anyone charged.
The exception is a public bid tabulation, where the agency publishes every bidder's unit price after opening. We know of exactly one letting in six years of Florida DOT records that carries floating dock and gangway pay items with all six bidders' numbers visible, and we have published the analysis of it and the underlying data. It is the obvious candidate for the first public pricing entry on this page. We would rather show an empty column than fill it with a number we cannot let you check.
In private pilots the pricing column is not empty, but those numbers belong to the customers whose files produced them and they are not ours to publish. That is a real limitation of this page and you should weigh it as one.
What we have got wrong
We keep a registry of every class of mistake the engine has made, with the guard that now prevents each one. It stands at 44 entries. Sorted by type, 17 are document reading, 7 are scope and counting, 7 are configuration and pricing policy, 5 are product classification, and the rest are process. Zero are arithmetic.
Two worth naming, because they are the kind of thing that never appears in a sales deck. The engine once read a dock as sitting on plastic tubs when the specification meant encapsulated foam, and quoted the job more than 40 percent low. Separately, it measured freeboard to the top of the float rather than the top of the deck, which is a difference of one frame depth, so it bought more flotation than the job needed on every dock while the totals still looked reasonable. Right money, wrong product.
Both are now gates that run before pricing. We publish them because a vendor with no list of failures either has not run enough real work or is not telling you about it.
Run this on any vendor, including us
Nothing in this protocol is proprietary and none of it is designed to favour us. If you are evaluating anyone in this category, take it and use it:
- Bring your ugliest recent package, not your cleanest. Include one where the specification and the drawings disagreed, and one where the customer expected you to design the layout.
- Keep your filed quotes out of the vendor's reach entirely. The leak usually happens by accident, through a quote PDF sitting in the same folder or a total mentioned in an email.
- Ask for reading accuracy and pricing accuracy as separate numbers, and refuse component accuracy as a substitute for either.
- Run one package twice on different days. The difference between the two runs is the variance, and it tells you how much any single answer can be trusted.
- Ask what the system does with a job unlike anything in your archive. If the answer is not that it tells you, keep looking.
- Agree the pass threshold in writing before anything runs.
If a vendor declines this format, that is information too, and it is cheaper to learn now than after a year of adoption.
Starting one
A pilot begins with a conversation about which product family to use and what the bar should be. It does not begin with a contract, and it does not require your live documents on day one. If you would rather not share anything yet, we will run a public bid package from your industry instead and you can watch what happens to it.
Email atishay@mavlon.co with the product family you have in mind, or book a call and bring the package you least want to read again.
Book a 30 minute call