What a Benchmark Actually Is on a Jobsite
A benchmark is nothing more than a number you trust enough to measure the next job against. It's the answer to the question every superintendent asks silently at the end of a bad week: "Is this normal, or are we getting killed?" Without a number, you're arguing about feelings. With one, you can point at a spreadsheet and say the drywall crew is hanging 40 percent less board per shift than they did on the last three towers, and now the conversation is about causes instead of opinions.
The trap with benchmarking software is that it gets sold as a dashboard full of gauges. Gauges don't build anything. What you want is a small set of honest numbers, collected the same way every time, that tell you where you're bleeding time and where you're actually good. Most firms already have this data buried in old daily reports and closed-out schedules. The work is pulling it out and being disciplined about how you compare.
The Handful of Metrics Worth Tracking
You can benchmark almost anything, which is exactly why most benchmarking programs collapse under their own weight. Pick a few metrics that drive decisions and ignore the rest. In twenty years I've never regretted tracking these:
- Duration per unit of work. Days per floor for a structure. Board per man-day for drywall. Linear feet of pipe per plumber per shift. Square feet of finished floor per day. Anything you can normalize to a repeatable unit so a 12-story job and a 22-story job can actually be compared.
- Percent Plan Complete (PPC). Of the tasks your crews committed to this week, what fraction actually finished as promised? This is the single most useful production number I know, and almost nobody outside Lean shops tracks it. A crew running 45 percent PPC isn't lazy — the plan is lying to them.
- Schedule variance at the milestone level. Not "are we behind overall," which is useless, but which specific milestones slipped and by how many days against the baseline.
- Reasons for non-completion. When a committed task doesn't finish, why? Prerequisite work wasn't ready, material didn't show, RFI open, crew pulled to another area, weather. This is the goldmine, and I'll come back to it.
Notice that only the first one is a productivity rate. The rest measure the reliability of your planning, which is where the real money hides. A crew that's fast but chronically blocked will lose to a slower crew that never waits.
PPC: The Benchmark That Actually Changes Behavior
If you take one thing from this article, make it PPC. At the end of each week, count the tasks your subs committed to in the weekly work plan, then count how many finished 100 percent complete — not 90, not "basically done," done. Divide. That's your PPC for the week.
The first time a team measures this honestly, the number is usually somewhere in the 50s. That's a shock to people who thought they were running a tight job, and it should be. A world-class team lives in the 80s and rarely touches 90-plus, because the last stretch is where the unpredictable stuff lives. If you're consistently above 90, you're probably sandbagging your commitments — the crews are only promising what's already guaranteed, which means the plan has slack you could be attacking.
The point of PPC isn't the grade. It's the paired discipline of writing down why every missed commitment failed. Stack up eight weeks of those reasons and a pattern jumps out of the page. Maybe 40 percent of your misses trace to material not being staged in time. That's not a scheduling problem, that's a procurement problem, and no amount of schedule pressure fixes it. This is the difference between benchmarking that produces a chart and benchmarking that produces a change order to how you run the job.
Building a Baseline You Can Trust
A benchmark is only as good as the data behind it, and construction data is filthy. Before you compare anything, you have to decide what counts. Does "start of drywall" mean the first sheet went up, or the day the crew was scheduled to mobilize, or the day the area was actually ready for them? Pick one definition and enforce it, because if half your projects measure mobilization and half measure first-board-hung, your "days per floor" number is garbage and you won't know it.
Three rules keep a baseline honest:
- Same start and stop events every time. Write the definitions down. "Floor complete" means inspected and signed, not "the crew says they're done."
- Normalize for the obvious differences. Board per man-day on a job with 20-foot ceilings and a hundred soffits is not comparable to a flat-ceiling apartment build. Note the conditions. A benchmark without context is a lie with a decimal point.
- Enough samples to mean something. One fast floor is luck. Five floors trending the same direction is a rate. Don't set a target off a single data point.
This is where good look-ahead scheduling software pays for itself quietly. When your weekly work plans, commitments, and actuals all live in one place instead of scattered across whiteboards and text messages, the historical record builds itself. Tools like LookAheadWall are built around location-based weekly plans and trade-flow sequences, so "days per floor per trade" isn't a research project — it's already in the structure of how you planned the work. You benchmark against your own history without having to reconstruct it from daily reports six months later.
Three Benchmarks Worth Comparing Against
There are really only three yardsticks, and each answers a different question.
Your own history is the most useful and the most honest. Same company, similar building types, similar crews and market. When this month's structure cycle is running 8 days a floor and your last three towers ran 6, that gap is real and it's yours to explain. Nobody can wave it off with "different market."
Organizational standards are the targets your company sets — the number estimating used to price the job. Beating the estimate means you're building float. Missing it means you're eating the schedule you sold. This comparison keeps the field and the estimators speaking the same language, which they otherwise never do.
Industry norms are the weakest yardstick, and I'd treat published "industry average" productivity figures with real suspicion. Conditions vary so wildly between a downtown high-rise and a garden-style multi-family that an external average tells you almost nothing about whether your specific crew on your specific job is doing well. Use industry numbers to sanity-check that you're not off by an order of magnitude, then get back to comparing against yourself.
How Benchmarking Goes Wrong
I've watched more benchmarking efforts die than succeed, and the failure modes rhyme.
Measuring so much that nobody looks at any of it. A forty-metric dashboard is a graveyard. Track five things people actually act on.
Turning the number into a weapon. The instant a foreman believes his PPC will be used to beat him in a meeting, the data goes bad. He'll only commit to sure things, stop reporting honest reasons for misses, and your goldmine dries up overnight. Benchmarks are for finding system problems, not for grading individuals. Protect that or the whole thing rots.
Comparing across conditions that aren't comparable. Winter versus summer. A repeat client with clean drawings versus a design-build fire drill. A benchmark that ignores context produces confident, wrong conclusions.
Chasing the metric instead of the outcome. If you reward high PPC, crews will hit high PPC by committing to almost nothing. The number goes up, the building goes nowhere. Always read a productivity number and a reliability number together — speed means nothing if the work has to come apart for a missed inspection.
Turning Numbers Into a Better Look-Ahead
Benchmarking is worthless if it doesn't change next week's plan. The loop is short and it runs every week: the crews commit to work in the weekly work plan, you measure what actually got done, you record why the misses happened, and you feed that straight back into how you build the next short-interval schedule.
Say your reason-code data shows that framing-to-rough-in is your recurring pinch point — the framers "finish" a floor, the trades pile in, and half the time something isn't actually ready. The fix isn't yelling. It's building a real buffer into the sequence: give frame-to-rough-in a 1 to 2 day gap for cleanup, punch, and inspection instead of stacking the next trade on the same afternoon. Now your benchmark has bought you a schedule change that will show up as higher PPC next month. That's the whole game. The number found the pattern; the look-ahead absorbed the fix.
A few durable rules of thumb from working this loop: never let a task sit on the weekly plan without a named crew and a made-ready check; treat any milestone slipping more than about 10 percent of its planned duration as a signal to look upstream, not to add overtime; and megger the runs and photograph the walls before anyone closes them, because a benchmark that includes rework is measuring the wrong thing.
Start Small and Stay Honest
You don't need a benchmarking program. You need one metric, tracked honestly, for one trade, for eight weeks. Pick PPC. Write down the misses and their reasons. At the end of two months you'll know more about why your job runs the way it does than any dashboard will ever tell you, and you'll have earned a real baseline to measure the next job against.
The firms that get good at this aren't the ones with the fanciest software. They're the ones who kept comparing against their own past, told the truth about the misses, and let the numbers quietly reshape how they plan the week. Do that long enough and "Is this normal, or are we getting killed?" stops being a feeling and becomes a question you can answer.