GTM Systems
Canon
Eleven spellings of one job title break lead scoring, segmentation and routing, and the usual fix rewrites values with no record of what changed or why.
- My role
- Owned the normalization engine, its confidence model, the review loop, and the audit I ran against all three
- Maturity
- Live across four surfaces, verified 2026-09-28. Engine suite verified 2026-10-01: 370 tests passing, 9 known defects each pinned by a failing test. Never shipped to a customer. Re-checked on HubSpot API changes.
- Decision
- Audit the product, not the documentation, before the first customer holds a copy they own.
- Boundary
- Auto-apply requires a confidence score at or above 85 and membership in the auto-acceptable field list. Title, persona and industry route to a person at any score.
4
Deploy Surfaces
370
Engine Tests Passing
9
Open Defects, Each Pinned
6 → 1
Wrong Unattended Writes
Built for Sidera. I designed and built all of it: the normalization engine, the review API, the in-CRM card, the database schema, the n8n orchestration, and the customer delivery process. I also commissioned the audit against my own work, wrote the brief that scoped it, and acted on the findings. That last part is in scope here because it is the part of this project I would most want to be judged on.
The operating problem
I pulled a contact export from a client portal and counted the job titles. "VP of Mktg," "V.P. Marketing," "Vice President - Marketing," and eight more variants describing the same seniority. Industry was worse: "SaaS/Technology," "Software," "Tech" and "Computer Software" in one column.
Lead scoring read all of it as mismatches, and everything downstream of those fields inherited the problem: segmentation, territory routing, ICP tiering.
The usual fix is a workflow that rewrites values on the way in. That works until it is wrong, and when it is wrong there is no record of what changed or why, and nothing to review.
The system decisions
Four choices shaped everything else, and each one closed off something easier.
Deterministic rules, not a model. "VP Mktg" has to resolve to Vice President of Marketing every time, not most of the time, because the output feeds lead scoring and territory routing. So the engine is a 38-entry abbreviation dictionary and priority-ordered pattern matching. Same input, same output, forever, and the reason is testable. A model would be faster to write and would resolve "Growth Hacker / Evangelist" with more imagination than I can, and it would also resolve the same title two different ways on two different days. The model is reserved for edge cases the rules cannot reach, with its confidence capped below the auto-apply threshold, so a guess can never apply itself.
The logic lives in Python, not in workflow nodes. Every part of this could have been built in n8n, and the first version of the idea was. The problem is that logic inside workflow nodes cannot be unit tested and cannot be moved. The engine is an importable package with its own suite, so the normalization decisions are portable and provable independent of any platform. n8n schedules it and moves data, which is what n8n is good at.
Two independent measures, not one score. Per-field confidence says how certain the engine is about one normalization. Record completeness says how much of the record exists at all. A contact can score 90 on seniority and still be an 80 out of 100 record because employee band is missing. Blending those into a single number produces something that sounds informative and decides nothing.
Canonical values go into a shadow layer. Every normalized value is written to a canon_* custom property, and HubSpot's native fields are never modified. This costs a property per field and some explaining. It buys the thing that makes the rest of the design possible: when Canon is wrong, and it is sometimes wrong, the original value is still sitting there untouched. A system that overwrites source data has to be right. This one only has to be recoverable.
Operating boundary
Confidence alone does not decide whether a change applies. Two conditions do.
| Tier | Rule | What happens |
|---|---|---|
| Auto-accept | Score at or above 85 and the field is auto-acceptable | Written to HubSpot directly, logged for audit |
| Human review | Score 50 to 84, or the field is always-review | Queued. Shows in the HubSpot card |
| Low confidence | Below 50 | Flagged, not suggested |
This is what makes two conditions matter rather than one. Run "VP Mktg" through the engine and seniority scores 90 and applies itself, while title scores 90 and still goes to a person. Same input, same score, different destination, because routing follows the field and not the number.
Why three fields never auto-apply
A confidence score says how sure the engine is about a value. It says nothing about what happens downstream if the engine is wrong. Those are separate questions, and a single threshold answers only the first.
So the second question gets its own list. title, persona and industry route to a person at any score, chosen by what an error does after it lands rather than by how often the engine gets them right.
Errors that propagate. Title feeds seniority, and seniority and department together feed persona. Industry feeds ICP tier, and tier feeds routing. A wrong title is not one wrong field, it is a wrong title and a wrong persona. A wrong industry becomes a wrong tier, which becomes a lead sitting in the wrong rep's queue.
Errors a person acts on directly. Persona is not an input to anything else. It is on the list because it decides what gets sent to someone. A wrong persona is a wrong email in front of a real prospect, and nobody traces that back to a normalization engine.
Against that, the auto-acceptable fields fail in place. A wrong region puts a contact in the wrong territory list. A wrong department costs one segment membership. Both are visible, both are correctable, and neither becomes something else on the way.
That division is a GTM question rather than an engineering one. Knowing which CRM fields are load-bearing means knowing what reads them three systems downstream, and no confidence curve contains that information.
Designed behavior, and where the numbers come from
Every number Canon acts on is a chosen value with a reason. None are fitted to a labelled dataset, and that is a position rather than an omission.
The threshold is 85, and the threshold is not the safety mechanism. Raising it to 95 would cut auto-accepts hard and still let a confidently wrong title through. The protection is structural instead, which is why the always-review list exists.
Derived values carry a discount, applied consistently. ICP tier scores at 0.9 of the industry score. Persona scores at 0.9 of the lower of seniority and department, since a persona inferred from two inputs cannot be more certain than the weaker one.
Record completeness is weighted by what each field unlocks: work email and domain 25, industry 20, seniority 15, department 10, lead source 10, less 10 when industry falls through to "Other."
What would change any of these is a labelled set of a few hundred records with independently recorded correct values. That measurement has not been run.
Evidence and verification
| Check | State |
|---|---|
| Engine suite | 370 tests passing, 2026-10-01. No skips, no external dependency, runs offline in under a second |
| Deployed surfaces | All four verified 2026-09-28: engine health endpoint, review API, database, and the card rendering inside a live HubSpot record |
| Audit coverage | 67 tests across 4 files written specifically to break Canon's documented claims |
| Open defects | 9, each pinned by a strict expected-failure test naming its finding |
| Measured error rate | On 100 seeded contacts, the scorers made 6 wrong unattended writes before the fixes and 1 after |
| Data model | 9 tables, 7 enums, 7 migrations, row-level security on every table in the live database |
Verification here is dated rather than counted. A test count says nothing about whether the thing it tested still behaves this morning, which is the mistake this project already made once.
The audit
Everything above is the system as designed, and it was all true before September 2026. So was this: Canon passed 244 tests and carried a critical defect in the workflow that touches the most records, and had done for months.
I commissioned an adversarial audit in September 2026, before the first customer. The brief was specific about one thing: check the product, not the description. Three earlier reviews had all asked a version of the same question, whether the description matched the code. One of them found six inaccurate claims and they were corrected. None of the three asked whether the code was right.
The method mattered more than the reviewer. Reading the code had already found nothing, three times. So the instrument was writing tests that assert what Canon claims about itself, then running them. The assertions failed, and that is how the defects appeared.
It returned seventeen findings. Three more surfaced afterward, and one open question resolved into a defect of its own.
What it found
The backfill workflow never called the routing gate. It read every normalized value the engine produced rather than the subset routed for automatic application. So it wrote canon_title, canon_persona and canon_industry, the three fields that are supposed to require a human at any confidence level. It wrote Unknown placeholders over real values. It sent no portal identifier, so nothing was recorded: no run, no review queue, no audit row.
It shipped marked active, and the setup guide tells a new customer to turn it on. So the first thing Canon would do in a customer's portal is run unreviewed across their entire existing database, before the reviewed real-time path ever saw a record.
Then the same shape again. A per-organization settings system, fully built and cached and loaded correctly, never reached the decision it configures. The caller resolved the organization fourteen lines after the routing decision had already been made, so every field routed against an organization literally named default. Customers could set a threshold. It would not apply.
Two defects, one shape: a mechanism built correctly, and a caller that skips it. Neither fails loudly. Each falls through to a plausible default.
A third looked like the same shape and was not. The approve route builds its HubSpot URL from a database enum whose values are singular, while every other call in the repository uses the plural form. I wrote it up as a defect: approvals would 404 and never reach the CRM.
Then I checked. One read-only call returned 200. HubSpot accepts the singular form, card approvals had been working the whole time, and I had inferred an API's behaviour from a naming convention and recorded the inference as a result. The write-up was corrected and the finding closed as cosmetic.
I am keeping that here because it is the same error the audit exists to catch, made by the person running the audit, and because the cost of checking was thirty seconds.
Why the timing was the whole thing
Canon is not hosted software. A customer receives their own repository, deployed once and then owned outright. That means a defect is not fixed centrally. Every customer holds a frozen copy of whatever shipped, and a bug arrives as an email rather than as a deploy.
Nobody had a copy yet. That is the only reason this is a case study rather than an incident report.
Decisions made under the audit
Unknown placeholders are never written. The broken backfill wrote Unknown for any field that scored zero. The easy fix was to keep writing them and treat the string as meaningful. I chose to write nothing instead. A placeholder overwrites real values, including values a reviewer had already approved, and Unknown counts as having the property, which defeated the backfill's own search for contacts still needing normalization. An absent field says "not known yet." A field containing Unknown says "known to be unknown," which was never true.
A failed audit write fails the request. The engine used to catch a Supabase failure, log a warning, and return 200 with the normalized values intact. That is why a broken database enum went unnoticed for months: n8n wrote to HubSpot as normal while nothing was recorded. It now returns 503. The cost is that a database outage stops normalization entirely. The alternative was worse: a 200 would stamp a record as normalized while permanently losing the review items that record generated. Degrading loudly beats degrading invisibly when the thing being lost is the audit trail.
The engine got its own secret rather than reusing the HubSpot token. They guard different doors: the HubSpot token authorizes writes to a customer's CRM, and the engine secret only proves a caller is ours. Reusing it would have put a CRM write token into one more place for no gain, and coupled two rotations that have nothing to do with each other.
What closed, and what has not
Twenty defects were found. Eleven are closed and nine remain, each still pinned by a failing test.
Closed. The backfill now routes through the gate and records what it writes. Four scoring defects are fixed, which moved the unattended error rate on a hundred seeded contacts from six wrong writes to one. Every write endpoint authenticates its caller, and the reviewer identity now comes from a value HubSpot signs rather than from the request body, so approvals record a real person instead of a blank. Every engine dependency is pinned to an exact version, after a deploy came up on a Python release the configuration file does not mention.
Not closed. The live database and the shipped migrations have drifted apart. A customer running the shipped SQL gets a different database than the one I validated against, which means my validation results do not transfer. Per-organization settings still do not reach the routing decision. The customer delivery process still points at my own infrastructure in four places.
And two scoring items stand. "Dir biz dev" normalizes to Engineering. One defect could not be measured at all, because the seed data uses only US email domains and that defect only fires on non-US consumer mailboxes. It reads as a clean result and is an absence of evidence.
What this demonstrates
Two hundred forty-four passing tests did not catch a critical defect in the most-used write path, because those tests were written by the same reasoning that wrote the code and inherited its assumptions. What caught it was writing tests that assert the documentation, then watching them fail.
The defects that mattered were never in the logic. They were in the wiring: a mechanism built correctly and a caller that bypasses it, falling through to a plausible default without an error. Two findings had that shape, including the critical one, and naming the pattern is what made the second one findable. A third candidate looked identical and dissolved under a thirty-second test, which is the more useful half of the lesson.
The design itself held. Always-review is enforced before the threshold is even read, and it held under every settings combination the audit could construct, including a score of 100 against a threshold of zero. The judgment that put three specific fields on that list is the part I would defend hardest, and it is the part a confidence score could never have produced.
And the fix surfaced a bug the audit could not have. Once the backfill writes only the routed subset, contacts sent for review never receive the property its search filters on, so it would return the same hundred records forever. That defect did not exist until the correct fix created it, which is an argument for a regression suite rather than a review.
Repository and technical artifacts
Python engine on Render, Next.js review API on Vercel, HubSpot app card via HubSpot Projects, PostgreSQL with row-level security on Supabase, n8n for orchestration and scheduling.
The engine is a licensed product delivered as source, so it stays private. The audit findings and the triage plan are the evidence for what this page claims and they contain no product logic. The regression suite stays private with the engine: tests assert input and output pairs, and enough of those reveal the dictionaries that took the longest to build.
Visible limitation
Canon normalizes the fields it was taught. A portal with a custom property nobody outside that company would recognize gets no help until someone writes the rules for it. The engine is extensible by design, but extending it is authoring work rather than configuration.
There is a second limit the audit made visible. Canon's error rate is only known against the data it has been tested on, and the seed set uses US email domains exclusively, so one scoring defect could not be sized at all. A rate measured on data that cannot trigger a defect is not a rate.