A builder looks at a bid and says, "This one will run over." He is usually right. He has priced hundreds of jobs. He knows the soil, the subs, and the clients who change their minds.
Then he hires an estimator. Then a second one. His instinct does not transfer. It was never written down. It was never tested.
That is the problem I build for.
Language models are not decision engines
The current wave of AI tools is built on language models. They draft, summarize, search, and answer. I use them in every system I build. They are good at language.
They are not built to answer the questions that run a small business. Which lead converts? Which job runs over budget? Which client is about to leave? Those are prediction problems. They have known structure, measurable error, and real costs when you get them wrong.
A chat window does not estimate a probability. It does not tell you how often it is wrong. It does not compare itself against your best estimator.
Dashboards fail differently. They describe the past. Averages, totals, bar charts. They show what happened. They stop short of what to do next.
The gap between the data and the decision is where most AI investments stall. That gap needs its own layer.
What a Decision Science Layer is
In the Dr. Data platform, the Decision Science Layer sits between the data and the decisions. It holds six things:
- Features. Clean, defined inputs built from the client's records. Missing values flagged. Sentinel and placeholder values caught before they distort a model.
- Models. From simple to complex, chosen by the data available.
- Evaluation. Held-out tests, calibration checks, and a comparison against the expert's own rules.
- A registry. Every model, its version, its training snapshot, its metrics, its date.
- Explanations. Every prediction shows why: coefficients, feature contributions, or the most similar past cases.
- Monitoring. Drift, freshness, and performance over time.
Data first. Decisions second. AI third. This layer is the bridge between the first two.
Models earn their place
Here is the rule every Dr. Data system follows. Every decision starts on the owner's confirmed rules. We capture those rules first, in a structured elicitation we call STZ. The owner confirms every one.
A model replaces a rule only when it beats that rule on data it has never seen. If it does not, the rules stay in charge, and the system says so.
This matters. A model that loses to the expert has no business making the call. A model that wins earns a place, and it keeps that place only while it keeps winning.
The model ladder
Small businesses do not start with ten thousand clean records. So we do not start with complex models. Each decision climbs a ladder as its data matures:
- Rules. The owner's confirmed logic. Always available.
- Describe. Profiles, distributions, percentiles, cohorts.
- Group. k-means segments and nearest-neighbor comparables, once there are enough complete records.
- Predict. Logistic and linear regression, once there are at least ten outcome events per feature, the common events-per-variable guideline in the statistics literature.
- Learn harder. Random forests and gradient boosting, only when they beat the regression on a held-out set.
- Perceive. Language and image models, for text and photos, trained and run locally.
The thresholds are firm defaults. A client's own data can move them.
Metrics that mean something
A model is only as useful as the evidence behind it. We report the metrics that answer real questions.
- AUC: can the model rank good leads above weak ones?
- Brier score and calibration: when it says 70 percent, does it happen about 70 percent of the time?
- Lift in the top decile: does the top tenth of the list convert better than the rest?
- MAE and interval coverage for estimates: how far off, and how often inside the stated range?
- Silhouette and stability for segments: are the groups real, or noise?
And one more, every time: performance against the owner's rules. If the model cannot beat the expert, it does not ship.
What it looks like in practice
Lead profiling. Public and business records come in: permits, licenses, parcels, property sales. Addresses are normalized. Records are joined and de-duplicated. Each prospect gets a profile. The owner's ideal-client rules score it first. As outcomes accumulate, logistic regression takes over, then gradient boosting if it earns the job.
Segments. For a real estate team, we clustered leads with k-means into audiences that each get their own outreach. The segments are checked for stability across resamples, so they hold up when the data shifts.
Estimating. Past jobs come in from the workbooks the builder already keeps. Nearest-neighbor search finds the most similar jobs. Regression estimates cost with a stated range. The estimator sees both, and decides.
Rule discovery. A decision tree trained on the decision log can surface rules the owner applies but never wrote down. Those rules go back to the owner for review. The owner decides whether they become part of the system.
Guardrails
Profiling demands discipline. Our systems use business and public records only. No protected traits are used as features or inferred. Opt-outs and suppression lists are honored in every campaign. Every data source carries a stated purpose. Every client's data stays in that client's own deployment.
How Dr. Data integrates it
Every Dr. Data system is one isolated platform per client: on the owner's machine, in a private cloud, or on their IT provider's server. Models train locally, on the client's data. Every prediction passes through a governed gateway and is logged with its inputs and its explanation. The owner sees every model on one page: what it does, how it scores, and whether it beats their rules.
The owner keeps the decision. The system gives them evidence.
That is what decision intelligence means to me. A system that knows how you decide, measures where the data can do better, and proves it before it acts.
Data first. Decisions second. AI third.
If you want to know where your business stands, start with our free Find Your Zone diagnostic.
Want AI built for your actual job? Book a discovery call.