AI Evals
AI evals are structured tests that measure how well an AI system performs on a defined task using examples with known answers. A construction firm can use them to test whether a tool reads its sub proposals or RFPs accurately before relying on it on bid day.
Why it matters in construction
Vendor demos often use clean proposals and well-organized RFPs. Your documents may include faxed plumbing quotes, handwritten alternates, or inconsistent scans. The only way to know how a tool handles them is to test it on them.
Evals make that test deliberate. They replace “it seems pretty good” with a record of what the tool found and missed across real proposals. The results help you decide which work needs review and whether the tool fits your process.
How it works
An eval needs a test set, a scoring method, and a report.
- A test set. Real documents from your own archive with the correct answers already established. For bid leveling, that is 20 to 50 past sub proposals and the leveling sheet the estimator built at the time. For RFP analysis, past RFPs and the key facts your team pulled from them.
- A scoring method. For each field, compare the model’s answer to the known answer. Exact match works for dates and totals. For exclusions and scope items, a person or a second model judges whether the extracted item matches the reference.
- A report. Record the accuracy for each field, every miss, and the document type that caused it.
Run the same eval on every tool you are considering. Run it again after a vendor updates its model or when you add a document type. Keep the misses. They show what an estimator needs to check by hand.
Do not rely on documents supplied by the vendor. Use your own, including the difficult ones.
Example in practice
Imagine a commercial GC choosing between two AI bid leveling tools. Instead of comparing demos, the chief estimator pulls 30 sub proposals from three recent projects, across drywall, electrical, and mechanical, along with the final leveling sheets.
Both tools process all 30. Tool A extracts totals correctly on 30 of 30 and catches 88 of 96 exclusions. Tool B gets 29 of 30 totals and 93 of 96 exclusions, but two of its three misses were on a scanned proposal where it silently returned nothing rather than flagging the page as unreadable.
The team picks Tool A. Its raw accuracy is lower, but it flags every miss as low confidence, giving an estimator a chance to catch it. Tool B silently returned a blank on a scanned page. The team saves the test set and reruns it each quarter.
Frequently asked questions
Do we need a data science team to run evals?
No. The most useful eval a GC can run is a folder of 20 to 30 real documents with the correct answers already known, run through the tool and checked by an estimator. That takes an afternoon and tells you more than any demo.
What is a good accuracy number for document extraction?
It depends on the field. Bid totals and due dates need a very high accuracy rate. Exclusion detection is harder. The tool should also flag uncertain results so an estimator can review them.
How often should evals be rerun?
Whenever the vendor ships a model update, whenever you change how you use the tool, and on a regular cadence like quarterly. Model behavior drifts, and the only way to notice is to measure.
Go deeper
- From the blog Best AI Preconstruction Software for GCs 2026 How general contractors can evaluate AI tools for takeoff, bid leveling, estimating, workforce planning, and forecasting, and when a connected platform wins.
- From the blog AI and Construction: The GC's No BS Guide to What Works in 2026 A founder's reality check on where AI actually delivers value in preconstruction, and where it still falls short.