CAT/TMS and Tooling Selection
The demo was flawless, which is exactly what worried Tomas. He was the head of localization for a medical-device company shipping into twenty-nine markets, and he had spent a Thursday afternoon watching a translation-management-system vendor click through a slide deck of ninety-one features: AI pre-translation with four engine connectors, a termbase with concept-level entries, an "AI Quality Score" that lit up green on every segment, automated workflows with conditional routing, a marketplace of plugins, single sign-on, and a dashboard so dense with gauges it looked like an aircraft cockpit. The account executive closed with the line every buyer hears: "It does everything." And it did. What the deck never showed, because Tomas never thought to ask, was whether the "AI Quality Score" was an MQM-aligned severity judgment or a repurposed edit-distance number wearing a quality costume; whether the tool could export a segment-level record of who post-edited what, against which engine draft, with which error scores, in a format an auditor could read; and whether, three years from now, when this company wanted to leave, it could take its translation memories, its termbases, and its quality history out in a standard format or would find them welded to a proprietary schema he would have to pay a consultant to liberate. Tomas bought the tool anyway, on the strength of the feature list and the green dashboard. Eighteen months later, during a regulator's audit of a mistranslated contraindication, he learned that the tool could not reconstruct which human had approved the segment, on what basis, or against what engine output, because it had never been built to record it. The feature list had ninety-one items. The one that mattered was not on it. This lesson is about how to select a computer-assisted-translation (CAT) tool and a translation-management system so that the thing that mattered to Tomas is on the list, at the top, and the feature-count theater is recognized for what it is.
The Tool Is the Operating Theatre, Not the Scalpel
Start with the two acronyms, because they get used interchangeably and the sloppiness hides the whole selection problem. A CAT tool (computer-assisted translation tool) is the segment-by-segment editing surface a linguist actually works in: the grid where the source sentence sits beside the target, where the machine-translation draft is pre-populated, where translation-memory matches and termbase hits surface in the sidebar, and where the post-editor makes the edit. A TMS (translation-management system) is the platform around that surface: it ingests files, splits them into projects and jobs, routes them to vendors and linguists, tracks status and deadlines, connects to the machine-translation engines, stores the linguistic assets, and hands the finished work back to whoever asked for it. In modern cloud platforms the two have largely merged into one system, and vendors sell them as one thing, which is precisely why buyers evaluate them as one undifferentiated blob of features instead of asking the harder question: what is this system actually for in an operation that owns its quality?
Here is the framing that reorganizes the entire decision. For a localization operation that has already gone MT-first, meaning machine translation (MT) or a large language model (LLM) pre-populates every segment before a human opens the file, the CAT tool and TMS are not where translation happens. Translation, in the sense of a human generating target prose from scratch, has largely stopped happening. What happens in the tool now is post-editing, meaning a human fixing and verifying a machine draft, and quality evaluation, meaning a human scoring that draft against an error typology to decide whether it ships. The tool is no longer the scalpel that does the cutting. It is the operating theatre: the environment where the machine draft is received, where the human intervention is performed, where the instruments are laid out, and, most importantly, where everything that happened is recorded so that if anyone ever asks what was done and why, there is an answer. A surgeon does not choose an operating theatre by counting how many lights are on the ceiling. They choose it by whether it supports the operation they perform and whether it keeps the record the hospital's lawyers will need. That is the shift in mindset this lesson is built on.
In an MT-first operation the CAT tool and TMS are not where translation happens. They are the operating theatre where a machine draft is received, a human intervention is performed, and, above all, a record of both is kept. Select for the operation and the record, not for the count of instruments on the tray.
Why Feature-List Theater Fools Good Buyers
Feature-list theater is the practice of winning an evaluation by breadth of capability rather than fitness for the buyer's actual workflow, and it fools good buyers for a structural reason: a feature list is a checklist, a checklist feels rigorous, and a longer checklist feels more rigorous than a shorter one. When Tomas compared three tools on a spreadsheet with feature columns, the tool with the most green cells won, and no column on that spreadsheet asked the questions that would have saved him. The theater works because the vendor controls the columns. Every capability the vendor built becomes a row; every capability the vendor lacks quietly never appears; and the buyer, who does not yet know what they will need in a regulator's audit two years out, cannot supply the missing rows themselves. The result is a selection optimized for demo impressiveness and comparison-spreadsheet completeness, which are the two things that correlate least with whether the tool will hold up the day something goes wrong.
The antidote is not a longer feature list. It is a shorter list of the right criteria, derived not from what vendors offer but from what an MT-first, quality-owned operation actually has to do: receive machine drafts from the engines it uses, let humans post-edit and score them against a real error typology, enforce approved terminology and clean translation-memory leverage, capture a segment-level record of provenance defensible to a client and an auditor, automate the movement of files without human packaging, and do all of this without welding the operation to one vendor forever. Six things. The rest of the ninety-one features are either in service of those six or they are decoration. The discipline of selection is refusing to score the decoration.
The Six Criteria That Actually Decide the Outcome
These are the criteria that determine whether the tool serves an operation that owns its quality, ordered roughly by how much damage getting them wrong will do. Each is defined so you can turn it into a demand you put to the vendor, not a box the vendor gets to check for you.
Criterion One: MT and LLM Integration You Actually Control
Every modern tool "integrates with AI," and the phrase means almost nothing until you interrogate it. The operation that owns its quality needs specific things from engine integration, and the demo will not volunteer whether they are present. First, multi-engine support: the ability to connect more than one MT or LLM engine and route content to the right one, because no single engine wins across every language pair and domain, and a tool locked to one engine has made a strategic decision on your behalf that you should be making yourself. Second, bring-your-own-engine and bring-your-own-key: the ability to plug in the specific engine you evaluated and chose, including a custom or fine-tuned engine, using your own API credentials, rather than being forced through the vendor's resold engine at the vendor's markup and the vendor's data terms. Third, glossary and translation-memory injection into the engine prompt or request, so the engine drafts against your approved terminology and your prior translations rather than its generic defaults, which is the difference between an engine that respects your assets and one that ignores them. Fourth, configurable pre-translation rules: control over which content gets machine-drafted, from which engine, at which risk tier, because in a risk-tiered operation some content must never be sent to a public engine at all, and the tool has to be able to enforce that boundary rather than blast everything to one endpoint.
The failure mode to fear is the tool that offers exactly one engine, resold, with no key control and no per-content routing. It demos beautifully, because one green "AI" checkbox looks identical to a real integration on a slide, and it quietly removes your ability to make the engine decision that the rest of your strategy depends on.
Criterion Two: Termbase and TM Handling as Enforceable Controls
A termbase is the store of approved terminology, the client's sanctioned word for each concept, with its forbidden alternatives and usage notes. A translation memory (TM) is the store of previously translated and approved segments, reused when a new source segment matches an old one exactly or fuzzily. Every tool has both. The question that separates them is whether these are enforceable controls or merely reference sidebars. A reference termbase shows the linguist the approved term and hopes they use it. An enforcing termbase runs a check that flags, or blocks, a segment where the source contains a term whose approved target is absent, and it injects that approved term into the engine's draft so the engine does not drift to a common synonym in the first place. The difference is the difference between a suggestion and a control, and in a regulated operation where an unapproved device name is a Major or Critical error, you need the control.
For translation memory, the discriminating questions are about cleanliness and leverage governance. Can the tool distinguish a human-approved TM entry from a raw machine-translated one, so that machine output does not silently pollute the memory that will pre-translate tomorrow's files? Can it apply penalties to fuzzy matches and to matches from lower-trust sources, so the leverage reflects real confidence rather than a naive percentage? Can it segment and align at the granularity your content needs? A TM that indiscriminately absorbs unverified machine output is not an asset that compounds quality; it is an infection vector that compounds error, and the tool's handling determines which one you get. Test both with your own real termbase and TM during evaluation, not the vendor's clean sample, because the vendor's sample is engineered to make the feature look effortless.
Criterion Three: QE and MQM Scoring Support, Not a Vanity Gauge
This is the criterion Tomas missed, and it is the one most disguised by theater. QE (quality estimation) is an automatic, reference-free confidence score a model assigns to a machine-translated segment, predicting how much editing it will need. MQM (Multidimensional Quality Metrics) is the human analytic evaluation framework, aligned with ISO 5060:2024, that categorizes each real error by dimension, accuracy, terminology, locale, fluency, and by severity, Critical, Major, or Minor, and computes a defensible quality verdict from them. These are different animals doing different jobs, and a tool must support both honestly, because conflating them is exactly the trap.
A tool with real QE support surfaces the automatic confidence score as a routing signal, letting you send low-confidence segments to fuller human review and higher-confidence ones to lighter checking, so human effort lands where it is needed. A tool with real MQM support gives the evaluator an interface to annotate a specific span with a specific error category and severity, accumulate those annotations across the file, apply the category weights, and produce a score with the rule that one Critical error fails the file regardless of how clean the rest reads. The vanity version, the version Tomas bought, is a single "AI Quality Score" that reads green and turns out to be an edit-distance or QE number rebadged as "quality," offering no error typology, no severity, no span-level annotation, and no way to produce the analytic verdict a standard requires. When you evaluate, ask the vendor to walk you through recording an actual Critical accuracy error, a flipped negation, on a specific segment, and watch whether the tool can categorize it, severity-rank it, and let that one error fail the file. If all it can do is show a green gauge, it cannot support the operation you are running.
An automatic quality-estimation score routes human effort; an MQM/ISO 5060 evaluation decides whether the file ships. A tool that offers a single green "AI Quality Score" and calls it quality is selling you a vanity gauge where you need a severity-scored verdict. Make the vendor record a Critical error live, or assume the capability is absent.
Criterion Four: Quality-Record and Provenance Capture
This is the criterion that fails silently until an audit, which is why it must be pulled to the front of the evaluation rather than discovered at the back. A quality record is the durable, exportable artifact that documents, for every segment, what happened to it: the source, the machine-translation draft and which engine produced it, the human edit, the terminology decisions, the error annotations and their severities, who did each step, and when. Provenance is the property that this chain is complete and reconstructable, that you can point at any shipped segment and show its full history rather than saying "a human looked at it." The revised ISO 18587, expanding its scope to cover AI and LLM "non-human translation output" and insisting the post-editor hold full professional-translator competence, and ISO 5060:2024, formalizing the severity-scored evaluation, both assume this record exists. An operation that owns its quality lives or dies on whether its tool captures it.
The discriminating questions are brutally concrete. Does the tool retain the machine draft as a distinct field alongside the final human target, or does the human edit overwrite the draft so the "before" is gone forever? Does it timestamp and attribute each action to a named user with a role? Does it store the MQM error annotations as structured data tied to segments, not as a free-text comment that evaporates on export? And, decisively, can it export all of this into a format a client or an auditor can read and a lawyer can rely on, or does the record live only inside the tool's proprietary interface, visible on screen but not extractable as evidence? Tomas's tool showed plenty on screen and exported almost none of it, which is the worst case: the appearance of a record with none of its evidentiary value. When you evaluate, demand an export of a fully worked, scored, post-edited file and read it as if you were the auditor. If you cannot reconstruct who did what to which segment against which engine draft, the tool does not capture provenance, no matter what the sales deck claims.
Criterion Five: Automation and API So the Loop Runs Itself
An MT-first operation at scale cannot have humans packaging files by hand, and the tool's automation and its API (application programming interface, the programmatic hooks that let other systems drive the tool) determine whether the loop runs itself or drowns your project managers in clerical work. The capabilities that matter: automated project creation and file routing based on content type and risk tier, so intake decides the workflow rather than a person; webhooks and API endpoints that let your content systems push source in and pull finished target out without a human export step, which is what makes continuous localization possible; connectors to the repositories and content platforms your organization actually uses; and automated pre-translation, quality-check, and gate steps that fire on every file rather than when someone remembers. The test is whether the tool can be driven headlessly by another system, because a tool that only works when a human clicks through its interface will become the bottleneck the moment your volume grows, and volume growth is the entire premise of going MT-first.
Beware the inverse theater here too: a rich API on the slide that is thinly documented, rate-limited into uselessness, or missing the specific endpoints your workflow needs. Ask for the API documentation during evaluation and have an engineer read it, because the gap between "we have an API" and "we have the endpoints your integration requires" is where automation projects die six months after purchase.
Criterion Six: Security and Data Governance
Security is last in this list only because it is usually the criterion buyers remember, not because it matters least; for regulated content it is often the gate that eliminates candidates before the others are even scored. The questions: Where is your data, and your clients' data, stored and processed, and does that satisfy the residency and privacy regimes your content falls under? When the tool sends a segment to an MT or LLM engine, does that segment leave your control, get retained by the engine provider, or get used to train their models, and can you turn that off with a contractual and technical guarantee rather than a checkbox? Does the tool support single sign-on, role-based access control, and audit logging of who accessed what? Does the vendor hold the security certifications your clients require, and will they sign the data-processing agreement your legal team needs? For an operation handling medical, legal, or financial content, a segment sent to a public engine that retains and trains on it can be a confidentiality breach independent of any translation error, and the tool's data-governance posture is what stands between your workflow and that breach. This criterion interacts with Criterion One: bring-your-own-key and per-content engine routing are not just engine-strategy features, they are the mechanism by which you keep high-liability content out of engines that would retain it.
Running an Evaluation That Resists the Theater
Knowing the criteria is half the job; the other half is running an evaluation process the vendor cannot steer back into a feature parade. The structure that resists the theater has four disciplines, and each one moves control from the vendor's slide deck to your operation's reality.
Discipline One: Write the Criteria Before You See a Demo
Author your weighted criteria, the six above plus any operation-specific ones, and their relative weights, before the first vendor demo, and freeze them. This is the single most protective move in the whole process, because the demo is engineered to make you weight whatever the vendor is best at, and a criteria set frozen in advance is immune to that reweighting. Assign weights that reflect real consequence: for a regulated operation, quality-record and provenance capture and QE/MQM support and security should dominate, while dashboard richness and plugin-marketplace size, the things demos lead with, should carry almost no weight or none at all. Write down the disqualifiers too, the capabilities whose absence eliminates a tool regardless of its other strengths, so that a tool which cannot export a provenance record is out before its lovely dashboard gets a chance to charm you.
Discipline Two: Evaluate With Your Content, Not Their Sample
Insist on a hands-on trial with your real files, your real termbase, your real translation memory, your real language pairs, and your real engine, driven by your real linguists doing their real workflow, not a guided tour of the vendor's pristine sample project. The vendor's sample is built to make every feature look effortless; your content is built to break things, which is exactly what you need to see. Have a linguist post-edit a genuinely hard file in the tool and report whether the termbase enforced, whether the engine injected the glossary, whether the QE routing was useful, and whether recording an MQM error was natural or a fight. Have an engineer test the API against a real integration slice. Have someone attempt the provenance export and read it as an auditor. Reality-testing with your own materials converts the abstract criteria into observed behavior, and observed behavior is the only evidence that resists a persuasive account executive.
Discipline Three: Make the Vendor Produce the Audit Artifact
Center the evaluation on one demand: produce, from a fully worked file, the exact artifact you would hand a regulator or a client's quality auditor, the segment-level quality record with engine provenance, human attribution, and severity-scored MQM annotations, exported in a readable, durable format. If the vendor can do this cleanly, the tool very likely has the deep capabilities underneath, because the record is downstream of MQM support, provenance capture, and export, so it cannot exist without them. If the vendor cannot, or produces a screenshot instead of an export, or a free-text blob instead of structured data, you have learned the most important thing about the tool in one request, and you learned it before signing rather than during an audit. This single demand does more to pierce the theater than any feature spreadsheet, because it is the one thing feature-count theater cannot fake.
Discipline Four: Score Total Cost of Ownership and Lock-In, Not Sticker Price
The sticker price, per-seat or per-word or per-project, is the smallest part of what a tool costs. The real cost includes the engine markup if you cannot bring your own key, the integration engineering to wire the tool into your systems, the training for linguists and project managers, the migration effort to get in, and, crucially, the migration effort to get out, which brings us to the criterion that quietly outweighs several of the others.
Vendor Lock-In and the Cost of Leaving
Vendor lock-in is the condition where the cost of leaving a tool is high enough to keep you using it even when it stops serving you, and in localization tooling the lock-in lives in your linguistic assets and your quality history. Your translation memories and termbases are the compounding value of years of human work; your quality records are the evidence of your conformance. If they can only leave the tool in a proprietary format, or cannot leave at all, then the tool owns them, and it owns you through them. Lock-in is not a hypothetical future annoyance. It is a present strategic risk that should carry real weight in the selection, because the tool you choose today you may need to leave in three years when your engine strategy changes, a client mandates a different platform, the vendor is acquired and the product degrades, or the pricing model turns hostile.
The Standard Formats That Are Your Exit Insurance
The insurance against lock-in is standard, portable formats, and a serious operation makes support for them a hard criterion. The load-bearing ones: TMX (Translation Memory eXchange), the standard interchange format for translation memories, so your TM can move to any other tool; TBX (TermBase eXchange), the standard for termbases, so your approved terminology travels; and XLIFF (XML Localization Interchange File Format), the standard bilingual working-file format, so your in-flight projects are not trapped. A tool that imports and exports clean TMX, TBX, and XLIFF is a tool you can leave, which paradoxically makes it a tool you can safely commit to. A tool that stores its assets only in a proprietary schema, or that exports a lossy, degraded version of them, has built a moat around your own data using your own data, and every year you use it the moat gets deeper because more of your work accumulates behind it.
The subtle failure to probe is lossy export: a tool that technically exports TMX but strips the metadata, the provenance, the match penalties, the custom fields, so what leaves is a hollowed-out shell of what you had. Test the round trip during evaluation: export your assets and quality records, examine what survived, and imagine rebuilding your operation from only what came out. What does not survive the export is what the tool has locked in, and a provenance record that cannot be exported in a durable, portable form is a provenance record you do not truly own.
Your translation memories, termbases, and quality records are the compounding value of your operation, and lock-in is the tool holding them hostage. Clean TMX, TBX, and XLIFF export is your exit insurance, and a tool you can leave is a tool you can safely commit to. Test the round trip; what does not survive the export is what the tool has locked in.
A Worked Tooling Selection
Walk a real selection the way it should go, because the worked example is where the criteria become a decision. Return to Tomas's operation, now doing the evaluation over again with the discipline he lacked the first time. The operation is MT-first, regulated (medical-device content across twenty-nine markets), quality-owned, and needs to be defensible under the revised ISO 18587 and ISO 5060. Three tools are in contention: Tool A, the feature-rich platform with the cockpit dashboard he originally bought; Tool B, a leaner cloud TMS with strong standards support; Tool C, an enterprise suite mid-way between them.
Step one: the frozen criteria and weights, written before any demo. Quality-record and provenance capture, weight 25. QE and MQM/ISO 5060 support, weight 20. Security and data governance, weight 20, with a disqualifier: any tool that cannot keep high-liability content off retaining public engines is out. MT/LLM integration with bring-your-own-key and per-content routing, weight 15. Termbase and TM handling as enforceable controls, weight 10. Automation and API, weight 10. Dashboard richness, plugin-marketplace size, and the other demo-candy: weight zero, explicitly, so they cannot creep back in. Lock-in is not a separate weighted row; it is a disqualifier woven through the record, TM, and termbase criteria via a hard requirement for clean, non-lossy TMX, TBX, and XLIFF export.
Step two: evaluation with real content. Each tool gets the same genuinely hard German contraindication file, the same real termbase and TM, the same chosen engine via bring-your-own-key where supported, and the same linguist. Tool A pre-translates through its resold engine only, with no key control, so the disqualifier on Criterion Six is already threatening: high-liability German content would flow through an engine whose retention terms Tomas cannot fully control. Tool B connects his own fine-tuned engine with his own key and injects the termbase into the request. Tool C supports bring-your-own-key but not per-content routing, so it can use the right engine but cannot cleanly wall off the tier that must never reach a public endpoint.
Step three: the audit-artifact demand. Each vendor must produce, from the worked file, the segment-level quality record with engine provenance, human attribution, and severity-scored MQM annotations, exported in a durable, readable format. Tool A produces a beautiful on-screen dashboard and a PDF that shows a green "AI Quality Score" and nothing reconstructable: no per-segment engine draft retained, no structured MQM annotations, no export that an auditor could trace, the exact failure that burned Tomas the first time, now surfaced before purchase instead of during a regulator's visit. Tool B produces a structured export retaining source, engine draft, human edit, named post-editor and evaluator, and MQM error annotations with severities, in a format that round-trips and reads cleanly, and its TMX/TBX/XLIFF export survives inspection with metadata intact. Tool C produces a partial record: it retains the engine draft and attribution but stores MQM annotations as semi-structured comments that degrade on export, and its TMX export drops the match penalties.
Step four: the scored result. Tool A, despite winning the original feature-count comparison decisively, now fails outright: it is disqualified on security (no per-content routing away from a retaining engine) and scores near zero on the two highest-weighted criteria, quality-record capture and MQM support, because its "quality" is a vanity gauge and its record does not export. The cockpit dashboard, weighted at zero, buys it nothing. Tool C scores respectably but its lossy export and missing per-content routing cost it the two disqualifier-adjacent criteria, and in a regulated operation those are not places to compromise. Tool B, the leaner tool that would have lost the original feature-count spreadsheet, wins on the weighted criteria that reflect the operation's real consequences, and, decisively, it is the tool Tomas can leave, because its standard-format exports are clean. The operation selects Tool B, and the sentence Tomas can now say to his regulator, his client, and his own leadership is the credential of this whole layer: here is the tool, chosen against weighted criteria we froze before any vendor spoke, that keeps our high-liability content off retaining engines, injects our approved terminology into the draft, lets our evaluators produce a severity-scored ISO 5060 verdict, captures an exportable segment-level provenance record an auditor can reconstruct, and holds our assets in standard formats we can carry out the day we choose to. Tool A's ninety-one features could not say one clause of that sentence.
Reading the Result the Way a Strategist Would
Step back and read what the disciplined selection actually bought, because it is the same dual posture this program keeps surfacing at every layer. Capability was not sacrificed: Tool B still drafts every segment with the engine, still runs the loop through its API, still leverages the TM and enforces the termbase, so the speed of the MT-first operation is fully intact. And quality was not left to a vendor's marketing: the criteria that decide whether the operation is defensible, provenance, severity scoring, security, portability, were weighted to dominate, tested against real content, and verified by an artifact the theater cannot fake. The strategist who runs the selection this way is not buying software; they are buying the operation's ability to prove its quality and to keep its assets free. That is a strategic decision wearing a procurement costume, and treating it as procurement, as Tomas did the first time, is how a regulated operation ends up unable to answer the one question that matters when a contraindication ships wrong.
Key Takeaways
- In an MT-first, quality-owned operation the CAT tool (the segment-level editing surface) and the TMS (the platform that routes, stores, and connects the engines) are the operating theatre, not the scalpel: the environment where a machine draft is received, a human post-edits and scores it, and, above all, a record of both is kept. Select for the operation and the record, not for the count of instruments on the tray.
- Feature-list theater wins evaluations by breadth because a longer checklist feels more rigorous, and the vendor controls the columns so their gaps never become rows. The antidote is not a longer list but a short list of the right criteria, derived from what the operation must do, with the discipline to refuse to score the decoration.
- Six criteria actually decide the outcome: MT/LLM integration you control (multi-engine, bring-your-own-key, glossary injection, per-content routing); termbase and TM handling as enforceable controls, not reference sidebars; QE and MQM/ISO 5060 support that produces a severity-scored verdict rather than a vanity "AI Quality Score"; quality-record and provenance capture that is exportable and reconstructable; automation and a real API so the loop runs itself; and security and data governance that keeps high-liability content off retaining engines.
- Quality estimation (an automatic reference-free confidence score) routes human effort; MQM (the human analytic evaluation aligned with ISO 5060, categorizing errors by dimension and by Critical/Major/Minor severity) decides whether the file ships. A tool that offers one green score and calls it quality is selling a vanity gauge where a severity-scored verdict is required. Make the vendor record a live Critical error, or assume the capability is absent.
- Quality-record and provenance capture (the exportable, reconstructable, per-segment history of source, engine draft, human edit, terminology decisions, attribution, and severity-scored annotations) is the criterion that fails silently until an audit, so pull it to the front. Demand an export of a fully worked file and read it as an auditor; if you cannot reconstruct who did what to which segment against which engine draft, the tool does not capture provenance whatever the deck claims.
- Vendor lock-in lives in your linguistic assets and quality history: if translation memories, termbases, and records can only leave in a proprietary or lossy format, the tool owns you through your own data. Clean, non-lossy TMX, TBX, and XLIFF export is your exit insurance, and a tool you can leave is one you can safely commit to. Test the round trip; what does not survive the export is what the tool has locked in.
- Run an evaluation the vendor cannot steer: write and freeze weighted criteria (with disqualifiers) before the first demo; evaluate with your real files, termbase, TM, engine, and linguists rather than the vendor's pristine sample; center the whole process on making the vendor produce the exact audit artifact a regulator would want; and score total cost of ownership and lock-in, not sticker price. The audit-artifact demand is the one thing feature-count theater cannot fake.
- The worked selection shows the payoff: the feature-rich tool that won the original feature-count spreadsheet is disqualified on security and scores near zero on the two highest-weighted criteria, while the leaner, standards-strong tool wins because it keeps high-liability content off retaining engines, produces an exportable severity-scored provenance record, and holds assets in portable formats. The credential of this layer is a sentence a ninety-one-feature deck cannot say: chosen against frozen criteria, defensible under ISO 18587 and 5060, and free to leave.
Skill.re