Introducing Keel Scan.
I built Keel Technologies because there was no simple way to know if an AI model was well-built.
The company did not happen overnight. For more than thirty years I have studied patterns, methodologies, and technologies across many fields, looking at how they work, when they hold, and what makes some frameworks durable and others fleeting. That long study is the foundation of Keel Technologies. The company has been decades in the making, and Keel Scan is the first product to come out of it.
I chose AI as the first field for two reasons. First, AI has real potential to benefit humanity. It can accelerate research, expand access to expertise, and make hard problems tractable. Second, AI right now is the Wild West, and I hope the tools Keel Technologies releases will help bring the industry into a more mature phase. A field cannot move as fast as it should without measurement, and measurement is what Keel Technologies is built to do.
You could try to measure models by running them on benchmarks. HellaSwag, MMLU, HumanEval, dozens of them, each measuring performance on a specific class of task. That tells you what the model can do. It does not tell you what the model is.
You could look at parameter count, layer count, context window. That tells you the shape of the model, but not the quality of how the shape was executed.
You could try to compare it to other models. But comparison depends on having a common ruler, and the AI field does not have one.
I spent a long time looking for a ruler like this. New models arrived every quarter, the field lit up with benchmark results, and I read those results and never came away knowing what any of the numbers meant about the models themselves. A model could score well on MMLU and poorly on HellaSwag. It could top a leaderboard and disappear the next quarter. It could be enormous and underperform something a fraction of its size. The numbers moved but the underlying question, is this model well-built, never got answered.
So I built the ruler.
Keel Scan is the ruler
Every Keel Scan report measures twelve things about a transformer model. Not what it does on any particular benchmark, but what the model itself looks like structurally. How the layers work together. Where compute is being spent. Where redundancy hides. How coherent the internal structure is. Where this model sits relative to every other model measured.
The output is a nineteen-page PDF. It contains a K# score, a single number on a zero-to-two-hundred scale, with one hundred as the current baseline across the ecosystem. It contains a letter grade, from A through F. It contains an architecture class, a coherence score, a thermal grade, a capability tier, and per-layer analysis of every one of the model's layers. It contains prescriptions, concrete actions ranked by impact.
What Keel Scan does not do
No training runs. No fine-tuning. No benchmark harness. A Keel Scan is a structural read on the model as it stands today, delivered in seventy-two hours.
Customers do three things with the report. They compare their model to others on the Keel Index, so they know where their work sits relative to the field. They act on the prescriptions, which are ranked by impact and often quantifiable in dollars per month. Most models have compute that is not doing useful work. And they use the K# score and grade in their own communications, giving external readers a calibrated reference.
The measurement technology is proprietary. Every claim on every report is calibrated against every other model we have measured. The numbers stand on their own.
The Keel Index
Every model that Keel Technologies scans goes on the Keel Index, a public reference of structural quality across the open-model ecosystem. As of today, thirty-four transformer models have been measured. Scores range from K# 83 to K# 153. The full Index sits on the Keel Scan product page, sortable, with sample reports for four representative models available for immediate download.
Availability
A Standard-tier Keel Scan is $1,500. The tier covers any dense transformer up to forty billion parameters hosted on Hugging Face. Larger models, gated repositories, and proprietary weights are supported at higher tiers.
Delivery for the Standard tier is seventy-two hours from commencement. Every scan produces the same nineteen-page report format. Full commercial terms are in the Service Schedule.
Why this matters
Better measurement means better decisions about which models to build on, which to retire, and where to spend the next month of engineering time. Multiplied across the field, better measurement means AI gets better sooner, and the benefits reach more people, more quickly. That is the point of the work.
Keel Scan is the first product from Keel Technologies. There will be more, in AI first, then in other fields where measurement is where value begins. The work behind Keel Scan applies broadly. When products in those other fields are ready to launch, they will follow.
For now, if you have a model you would like to see measured, the intake form is on the Keel Scan product page. I look forward to seeing what you have built.
Through measurement, we grow intelligence.
-David