How to Capture Estimator Knowledge Before It Walks Out

I once sat with an estimator who had been pricing floating docks for over fifteen years. I asked him to explain how he decided the float family for a particular job. He pulled up the spec, pointed to a paragraph, and said "this one is foam." I asked how he knew. He said "because it says marine environment and the designer referenced the Army Corps guidance, and whenever you see that combination it means encapsulated foam, not polyethylene tubs." I searched the document. Neither the words "encapsulated foam" nor "polyethylene" appeared anywhere. He was right. The correct float family was derivable from two indirect signals that the spec writer probably did not realize they were sending. Knowing how to capture estimator knowledge starts with understanding that the knowledge looks like this: invisible patterns built up over years, applied in seconds, and impossible to find unless someone who has them sits with you and shows you what they see.
This guide is about the problem that every custom manufacturer quietly worries about: what happens the day your best estimator is no longer available. It is not a product pitch. The principles apply whether you use software, hire a successor, or try to do it with documentation and training. We happen to build the software version, and we have the scars from learning how hard the knowledge-capture part really is. That experience is what this guide comes from.
The companion piece, why quoting is hard to automate, covers the structural reasons quoting has no infrastructure. This one is about the knowledge itself: where it lives, why standard documentation fails to capture it, and the approach that actually works.
Your best estimator is a single point of failure
Every custom manufacturer has a person who knows the pricing. Sometimes it is one person. Sometimes it is two or three. In either case, the ratio of knowledge-holders to company revenue is extreme. A firm doing $30 million a year in custom fabrication may have its entire pricing capability concentrated in one or two heads. The estimation team's combined experience represents decades of accumulated judgment, and nearly all of it is undocumented.
This is not a technology problem. It is a business continuity problem dressed up as a staffing one. When the VP of Sales says "we need to hire another estimator," what they usually mean is "we need another person who can do what Dave does." But what Dave does is not a teachable procedure. It is a collection of conventions, relationships, interpretation shortcuts, and market intuitions that took fifteen years to build. A new hire can learn the steps in a month. The judgment takes years.
And the window is short. The average age of a senior estimator in specialty fabrication is well into the fifties. Retirement is not a hypothetical for most shops. It is a five-year horizon, sometimes less. The question is not whether this knowledge will walk out the door. It is whether anything will be left behind when it does.
What happens when the knowledge leaves
The company does not just get slower. It gets wrong. And the wrongness is hard to diagnose because nobody knows what "right" looked like from the inside.
The symptoms show up gradually. More bids lost without a clear reason. Jobs won at margins that turn out to be thinner than expected, because a rate correction the senior estimator applied from memory is now missing. Change orders in the first few months of a project because a spec interpretation convention was never transferred: the new estimator read the specification literally and priced the named brand, when the previous estimator knew from experience that the engineer accepts the generic.
The most painful symptom: the new estimator's quotes start coming back "off," and nobody can tell them why. The senior estimator would have said "you priced the gangway with pedestrian loading, but this district always specs vehicular loading for maintenance access even on pedestrian-only structures." That knowledge is now gone. The new estimator prices pedestrian loading because that is what the document says, and the shop eats the difference when the submittal comes back rejected.
None of this shows up in an exit interview. "Can you document your process?" produces a list of steps, not the knowledge. "Can you train the new person?" produces a few weeks of shoulder-surfing that covers the easy cases and misses the hard ones, because the hard ones are hard precisely because they depend on pattern recognition that the senior estimator applies unconsciously.
How to capture estimator knowledge: why standard documentation fails
The instinct is to document the process. Write an SOP. Record the steps. Create a training manual. This is what operations teams do, and for manufacturing processes with repeatable steps, it works.
For quoting, it fails, and understanding why it fails is the key to understanding how to capture estimator knowledge for real.
An SOP captures sequence: step one, open the package. Step two, read the specification. Step three, check the drawings. Step four, price the materials. Step five, calculate labor. Step six, add margin. Step seven, review and send. Every estimator follows roughly this sequence, and writing it down teaches nobody anything they did not already know.
The knowledge lives not in the sequence but in the decisions within each step. "Read the specification" is a step. "Notice that the spec references the Army Corps coastal guidance, which means the float family should be encapsulated foam even though the spec never says 'foam'" is knowledge. "Price the materials" is a step. "Use the backup vendor for 5086 sheet because the primary vendor's lead time went to fourteen weeks last month and you will miss the bid deadline" is knowledge.
You cannot capture the decisions in an SOP because the decisions change with every project. A floating dock quoted for a protected harbor in the Great Lakes is a different exercise from the same dock quoted for an exposed Gulf Coast site, even though the SOP steps are identical. The knowledge is in knowing which decisions differ and why.
We tried the documentation approach ourselves when we started building a quoting engine. We asked an experienced estimator to describe their process in detail. We got a clear, well-organized explanation that covered about thirty percent of what they actually do. The other seventy percent only emerged when we sat with their past quotes and asked "why did you do this?" on specific lines, in specific projects, with specific numbers in front of both of us.
The five places quoting knowledge actually lives
If you want to understand how to capture estimator knowledge, you first need to know where it is. Through our engine builds, we have identified five locations. None of them are labeled "estimating knowledge." None of them are searchable. All of them have to be mined actively.
1. The Excel formulas
Not the values. The formulas. This is the single most important lesson of any knowledge-capture effort, and almost everyone misses it the first time.
Open a costing workbook in normal mode and you see numbers: a unit price of $84, a total of $12,600, a margin of eighteen percent. Open it in formula mode and you see the logic: this cell multiplies that cell by 1.15, that cell pulls from a rate table on a hidden sheet, this one uses a MAX function to pick the higher of two vendor quotes as a safety margin. The formulas encode decision rules that the estimator set up once and never explained to anyone.
We missed a cost worth tens of thousands of dollars because it was stacked into a formula that looked like a straightforward rate cell in values mode. The cell showed $230. The formula showed that $230 was actually a base rate plus a second structural member factored into the same cell. The estimator had added the second member to the formula five years earlier when they learned that a particular design always doubled that component. They never told anyone. Why would they? The spreadsheet handled it.
The lesson: every workbook review must read formulas, not values. Export the workbook with formulas visible. Trace every cross-reference. Look at hidden sheets. Check named ranges. The logic in the formulas is the closest thing to documentation most estimating teams will ever produce, and nobody treats it as documentation because it looks like a spreadsheet.
2. The rate corrections from memory
Material prices change. Vendor minimums shift. A supplier the shop has used for ten years starts slipping on quality, so the estimator quietly switches to the backup. None of these corrections appear in the workbook. The workbook has the base rate. The estimator adjusts it in their head or on a scratch pad, types the adjusted number into the quote, and the adjustment disappears.
When we started comparing engine-generated quotes against filed quotes, the first class of discrepancy we could not explain was rate corrections. The engine would produce $185 per linear foot for a guardrail. The filed quote said $200. The difference was that the estimator knew the base rate from six months ago was stale, the current rate was roughly eight percent higher, and they had rounded up to a clean number. None of that was written anywhere. It was knowledge that existed for the thirty seconds between looking up the old rate and typing the new one.
Capturing rate corrections requires comparing filed quotes to the workbook formulas over time. The gap between what the formula says and what the estimator actually typed is the correction. Track these across twenty or thirty jobs and patterns emerge: the estimator consistently adds a percentage for certain materials, consistently rounds up for certain vendors, consistently uses a different rate for rush jobs. These are mineable conventions, but only if you have the paired data.
3. The vendor relationships
Minimum order sizes. Real prices versus list prices. Lead times that matter for liquidated damages. Which vendors actually deliver on schedule and which ones you pad by two weeks. This is competitive intelligence, and it lives in the estimator's phone contacts, email threads, and memory.
A new estimator looking at the same product specification will use the published catalog price. The senior estimator knows that for orders over a certain volume, the real price is thirty percent below list, and that you have to call, not email, to get it. That single piece of knowledge can be a five-figure difference on a mid-size project.
Vendor knowledge is the hardest category to mine because it is inherently relational and time-sensitive. Minimums change. Contacts move. Pricing agreements expire. The best approach we have found is to capture vendor-related decisions at the quote level: for each subcontracted or purchased line item, record which vendor was selected, what price was used, and whether it differed from the published rate. Over enough quotes, the vendor-selection logic reveals itself.
4. The spec interpretation conventions
When the spec says "or approved equal," does the engineer mean it? When a general note says "all welding per AWS D1.1," does the shop MIG weld everything or does the estimator know that this district's inspector rejects MIG on structural connections? When the addendum changes a dimension on one drawing but not the corresponding detail on another, which one governs?
These conventions are district-specific, engineer-specific, and sometimes project-specific. The senior estimator has built a mental database of which consulting engineers are strict, which are flexible, which public agencies run a tight ship, and which ones will entertain a value-engineering alternate. This database is never documented because nobody thinks of it as data. It is just "knowing the players."
We mine spec interpretation conventions by comparing the estimator's quote against a literal reading of the specification. Every place where the estimate departs from the literal reading is a convention. "The spec says brand X hardware but the estimate uses generic" is a convention. "The spec says painting per SSPC but the estimate includes hot-dip galvanizing instead" is a convention. Once identified, each convention needs a stated mechanism: why does this departure work? The answer is usually "the estimator knows this engineer" or "this is standard practice for this district," and that explanation becomes the rule.
5. The judgment calls
That general contractor always submits change orders in the first month. That owner representative will hold the retainage past the contractual period. This project site has access constraints that add thirty percent to installation labor. These are judgment calls that the estimator folds into the price as contingencies, risk factors, and line-item adjustments. They are the least mineable form of knowledge because they change with every project and depend on information that may not appear in any document.
Judgment is the one category that cannot be fully captured, and being honest about this is important. When someone asks how to capture estimator knowledge, the honest answer is: you can capture about eighty percent of it through the methods above. The remaining twenty percent is genuine human judgment that varies with every project. The goal is not to replace the judgment. It is to make sure the other eighty percent does not vanish when the estimator who holds it is no longer there.
The corpus-mining approach
The approach that actually works is not documentation. It is mining. You do not ask the estimator to describe their knowledge. You look at what they have done and reverse-engineer the rules from the evidence.
The evidence is the paired set of bid packages and filed quotes. Each pair is a case study: here is what was asked, and here is how the estimator answered. The gap between a literal reading of the package and the estimator's answer is where the knowledge lives. Mine enough pairs and patterns emerge. Those patterns are the rules.
Here is how it works in practice.
Step one: collect pairs. Gather as many past jobs as possible where you have both the bid package (drawings, specifications, addenda) and the company's own answer (filed quote, costing workbook, or internal pricing sheet). One without the other is useless. Aim for at least twenty pairs to start. More is better. Organize them by product family so patterns within a family are not diluted by patterns from a different family.
Step two: reverse-engineer the math. Before you can mine conventions, you have to understand the costing identity. How does this company compute a price? What is the margin structure? How do they mark up subcontracted versus in-house work? What are the rate eras, the time windows where a given material rate was in effect? Verify your understanding by reproducing the estimator's filed totals from their own workbooks. Not within five percent. To the cent. If you cannot reproduce their arithmetic, you do not yet understand their process, and everything downstream will be noise.
Step three: mine the conventions. Go through the pairs and look for repeating decisions. Every convention that makes it into the rule set must clear a bar: at least three independent projects showing the same pattern, at least an eighty percent fit rate across the corpus, and a stated mechanism explaining why it works. Below the bar, it is a flag, a bracket showing the range of past decisions, or a question to ask the estimator. Never a silent default. The purpose of the bar is to separate knowledge from coincidence.
Step four: validate with the estimator. This is where the approach differs from pure data analysis. You bring the mined conventions back to the estimator and ask: "Is this right? Is this what you actually do, and if so, why?" About seventy percent of the time, they confirm it and add context you could not have guessed. Twenty percent of the time, they correct it: "No, I only do that for the coastal district, not the inland ones." Ten percent of the time, they are surprised: "I did not realize I was doing that." All three responses are valuable. The corrections prevent false rules. The surprises are often the deepest knowledge, the patterns the estimator applies unconsciously.
Step five: test the rules. Apply the mined rules to the held-out blind set, the jobs the estimator quoted that were never included in the mining. If the rules reproduce the estimator's decisions on jobs they were not trained on, the knowledge has been captured. If they do not, the rules are overfit to the training set and need refinement. This is the quoting equivalent of a test suite, and it is described in detail in our pilot protocol.
Three kinds of estimator knowledge
Not all knowledge is equally capturable. Understanding the categories helps set realistic expectations.
| Kind | What it looks like | How to capture it | Example |
|---|---|---|---|
| Derivable | Follows from published data or physical law. | Build the calculation once. Verify it. | Buoyancy calculation for a floating dock, ADA gangway slope for a given tide range. A formula handles it. |
| Mineable | Hidden in past quotes, consistent across projects, has a mechanism. | Corpus mining with the five steps above. | The estimator always uses the backup vendor for 5086 sheet on rush jobs. The estimator adds ten percent to stainless hardware for this district. |
| Only-the-estimator-knows | Project-specific judgment. Changes every time. | Cannot be captured as a rule. Make it a question the system asks at quote time. | Whether the owner representative will hold retainage. Whether the site has access constraints. Whether to bid the alternate. |
The goal is not to eliminate the third category. It is to make sure the first two categories survive when the person who currently holds them is no longer there. Derivable knowledge should be calculated, not remembered. Mineable knowledge should be encoded as rules, not carried in someone's head. Only-the-estimator-knows items should be framed as explicit questions with brackets showing the range of past answers, so that a new estimator knows the question exists even if they do not yet know the answer.
The proportions vary by company and product family. In our experience, derivable knowledge accounts for roughly twenty percent of a quote's value, mineable knowledge for about sixty percent, and judgment for the remaining twenty. Those percentages mean that a well-executed mining effort can preserve eighty percent of the estimator's knowledge. That is the difference between the company getting wrong when the senior estimator leaves and the company getting slower but staying accurate.
What the estimator actually needs to contribute
The biggest fear estimators have about knowledge-capture efforts is that they are being asked to replace themselves. It is worth addressing this directly, because their cooperation is essential and their resistance is rational.
The estimator's contribution is about two hours per week for the duration of the mining effort. What those hours look like: reviewing the mined conventions and saying "yes, that is right" or "no, I only do that for these cases." Answering questions about specific line items in past quotes. Explaining why a particular workbook formula does what it does. Confirming or correcting the rate-era timeline.
What it does not look like: sitting in a room dictating their process into a recorder. Writing a manual. Training a replacement in real time. Those approaches ask the estimator to articulate knowledge they apply unconsciously, which is like asking a native speaker to explain the grammar rules they follow: they can do it, but the explanation covers about a third of the actual behavior, and the rest comes out as "it just sounds right."
The mining approach is different because it starts with the evidence, not the explanation. The evidence is the estimator's own past work. The questions are specific: "On this job, your filed quote for the guardrail was $200 per linear foot, but the formula gives $185. What was the additional $15?" That question has a specific, verifiable answer, not a general description of a process. It is much easier for an estimator to answer "what did you do here?" than "how do you estimate?"
And the result is different too. The estimator does not become redundant. They become the reviewer instead of the sole producer. The mined rules handle the eighty percent that is repeatable. The estimator handles the twenty percent that is judgment. Their throughput increases because they are no longer re-deriving rules they already know on every job. They are checking whether the rules applied correctly and making the calls that only a human can make.
How long it takes and what to expect
Collecting twenty-plus pairs: one to two weeks, mostly calendar time waiting for the company to dig out old bid packages and match them to quotes.
Reverse-engineering the costing identity: one week of concentrated work. This is the phase that feels the most tedious and is the most important. Get the math wrong here and everything downstream is noise.
Mining conventions across the corpus: two to four weeks, depending on how many product families the shop prices. A company that only makes floating docks is faster. A company that prices docks, gangways, fixed piers, boat ramps, and bulkheads has five families to mine, each with its own conventions.
Validating with the estimator: ongoing, two hours per week. This runs in parallel with the mining and continues through the blind test.
Blind testing: one to two weeks. Running ten held-out jobs, scoring them, diagnosing failures, and deciding which rules need refinement.
Total elapsed time: six to ten weeks from starting the collection to having a validated set of rules. The effort is front-loaded. Once the corpus is mined and the rules are validated, maintaining them is a small ongoing investment: updating rate eras, adding new conventions as the shop evolves, and re-running the blind test periodically to confirm accuracy has not drifted.
The output is not a document. It is a tested rule set, a rate-era timeline, a library of spec-interpretation conventions, and a blind-test baseline. Together, these form the quoting corpus: the company's pricing knowledge in a form that survives the departure of any individual.
How to capture estimator knowledge: a starting point
If you have read this far and the problem resonates, here is where to start, regardless of whether you ever buy software.
First, pick three recent jobs where you have both the bid package and the filed quote. Open the costing workbook in formula mode. For every cell that contributes to the total, write down what the formula does, not what the number is. This exercise alone will surface two or three things nobody on the team knew the workbook was doing.
Second, compare the filed quote to a literal reading of the specification. Every line where the estimate departs from the literal spec is a convention. Write each one down with the estimator's explanation of why the departure works. You are looking for "this engineer accepts generic hardware" and "we always add ten percent on stainless for this district" and "piling is by others even though the spec is ambiguous, because the marine contractor always takes piling." These are the rules that a new estimator would not know and cannot find in the documents.
Third, ask the estimator about rate corrections. For the three jobs you selected, compare the base rates in the workbook to the numbers that actually appeared in the quote. Where they differ, ask why. The answers will fall into rate-era effects (the price changed), vendor effects (a different supplier was used), and judgment effects (a contingency was added). Sorting these is the first step toward a rate-era timeline.
These three exercises will take about a day of the estimator's time and produce more usable knowledge than a month of process documentation. They work because they start with the evidence, not the narrative. The estimator does not have to describe their process in the abstract. They just have to explain what they did, on a specific job, with specific numbers in front of them.
If you want to see what a full-scale version of this looks like, including the blind test, the failure registry, and the pass bar, we publish the protocol at mavlon.co/pilot. The structural reasons why all of this is necessary are in why quoting is hard to automate. And if your quoting work involves marine structures, the nine specific interpretation decisions that trip up every dock and gangway quote are in quoting custom gangways and docks.