When I design a decision system for a small business, the first thing I map is who makes each decision today. Usually it is one person, an estimator or an owner, who carries the rules in their head and applies them dozens of times a week. My work starts by writing those rules down with them, one at a time, and then building a system that runs on those rules from the first day.
People often expect a large language model to sit at the center of that system. In my architecture, the center belongs to the expert's rules and to the machine learning models that earn their place by beating those rules on held-out data. Language models work around the decision, handling the reading, searching, and writing that surround it.
I design every system in three lanes. The first lane is code: rules, parsers, and classic machine learning such as logistic regression, k-means clustering, and gradient boosting. Every decision, number, date, and document classification runs in this lane, where each output traces back to a rule or a model version I can show the client. The second lane holds open-weight language models running on hardware the client or their IT partner controls. The third lane is a paid frontier model on the client's own key, reserved for tasks where an open model came up short on a measured test.
The second lane carries a large share of the daily work. It turns document text into embeddings, so a question lands on the right page. It reranks search results by relevance. It drafts replies, follow-ups, and briefs, and it pulls free-text fields like a scope description once the deterministic parsers have handled the structured fields. Each output carries its source, and anything that leaves the building waits for a person to approve it.
Keeping language models in this supporting role makes the whole system easier to trust. When a quote comes out high, the expert can see the rule that set the margin. When a lead ranks first, the score points to the features behind it. The language model helps the expert find, read, and write faster, while the decision stays explainable from end to end.
This design also keeps daily costs steady. Open-weight models like nomic-embed-text for embeddings and qwen2.5 for reranking run locally through Ollama, so a busy week of searching and drafting runs on a fixed hardware cost instead of per-call charges. Each model is pinned to a version, its license is recorded, and every call passes through a gateway that checks the user's role, the data's class, and the cost.
When a business owner asks me where the AI lives in their system, I give them a two-part answer. The intelligence lives in their own rules and in the models that proved themselves against those rules. The language models are the assistants that make that intelligence easy to use every single day.
Common questions
Should a large language model make business decisions?
In our systems, decisions run on the expert's confirmed rules and on machine learning models that beat those rules on held-out data. Language models support the work around each decision, such as search, summaries, and drafts, so every decision traces back to a rule or a model version.
How do you keep AI decisions explainable?
Every decision, number, and classification runs in code, where each output traces to a rule or a pinned model version. When a quote is high or a lead ranks first, the expert can see the rule or the features behind it.
Related reading
- Why Every Managed Intelligence Platform Needs a Decision Science Layer
- Find Your Zone: a free three-minute diagnostic
- Why I Run Daily AI Work on Open-Weight Models
- Why the Language Model Comes Last in Our Document Pipeline
Want AI built for your actual job? Book a discovery call.