The AI-Supported Eligibility Workflow
The eligibility unit at a county human-services office runs on a clock that never stops. A worker there carries a queue of 200 active cases and applications in various stages, and the SNAP (Supplemental Nutrition Assistance Program, the federal food-assistance benefit once called food stamps) processing standard gives her thirty days from application to determination, seven days for the expedited cases where a household has almost no income and almost no money in the bank. On a Monday in March she opens an application from a single mother of three who marked the expedited box. The file is forty pages: pay stubs, a lease, a utility shutoff notice, a letter about a child receiving Supplemental Security Income (SSI), and a handwritten note explaining that the second job ended in February. The worker has roughly eleven minutes per case if she is going to clear the day's queue. She opens the agency's AI eligibility-support tool, feeds it the documents, and it returns a clean income calculation, a flagged categorical-eligibility pathway, and a draft determination in ninety seconds. The ninety seconds is the easy part. The next several minutes, the part this lesson is about, are where the family's food for the month is actually decided, and where a workflow either protects that family or quietly fails it.
Why Eligibility Is the Hardest Place to Get This Right
Eligibility work looks, from the outside, like the most automatable thing in human services. It is rule-bound. There are income thresholds, asset limits, household-composition tests, and deduction schedules. It feels like arithmetic, and arithmetic is what computers do. That appearance is exactly what makes it dangerous to hand to a model carelessly.
The reality is that benefits policy is not arithmetic. It is a dense, layered body of federal regulation, state options, and local administrative practice that changes every year and contains nested exceptions that swallow the general rules. The SNAP gross-income test, set at 130 percent of the federal poverty level, looks like a simple threshold until you remember that a household with a member receiving SSI or Temporary Assistance for Needy Families (TANF, the cash-assistance program) may be categorically eligible, which means the gross-income test does not apply to them at all. A determination that runs the income test and stops has misapplied the policy even though every number in it is correct. The worker in our opening scene was looking at exactly this situation: a child on SSI changed which test governed the case.
This is the core tension of the AI-supported eligibility workflow. The model is genuinely good at the mechanical parts: reading a pay stub, adding income, applying a standard deduction, checking a number against a threshold. It is unreliable at the parts that require knowing which rule governs, whether that rule is current, and whether the document in front of it actually says what it appears to say. And the consequence of getting it wrong is not abstract. A wrong denial means a family with three children does not have food assistance this month. A wrong approval that is later reversed can generate an overpayment claim that follows a family for years. The stakes sit on both sides of the determination.
There is a second reason eligibility resists clean automation, and it is one that workers feel daily even if they have never named it. The documents people submit do not arrive in the tidy form a rule expects. A pay stub may show a gross figure, a year-to-date figure, overtime, a one-time bonus, and a pre-tax deduction, and the policy question of which of those counts as countable income for this program is not something the document announces. A self-employed applicant may submit a shoebox of receipts and a handwritten ledger, and the income calculation requires judgment about allowable business expenses that no extraction tool can make. A household may include a relative whose presence is ambiguous: are they a member of the household for SNAP purposes, which turns on whether they purchase and prepare food together, a question the documents rarely answer directly. The model can read what is on the page. It cannot interview the family, and much of eligibility turns on facts that only a conversation surfaces. A workflow that forgets this and treats the documents as the whole truth will produce confident determinations built on an incomplete picture.
Eligibility feels like arithmetic and is actually policy. The numbers are the easy part. Knowing which rule governs is the part that decides the case, and it is the part the model is least reliable at.
The Five Stages of the Workflow
An AI-supported eligibility workflow that protects people has five distinct stages, and the order matters. Each stage has a clear owner, and exactly one stage, the determination, is reserved for the human and cannot be delegated under any circumstance. Think of the workflow as a relay where the model runs the first leg and the last leg is run by a person who is fully accountable for crossing the line.
Stage One: Intake and Document Extraction
The first stage is reading the file. The applicant has submitted documents, often a chaotic mix of phone photos, scanned pay stubs, a lease, benefit letters, and handwritten explanations. The model's job here is extraction: pull the gross wages from each pay stub, identify the pay frequency, find the household members, locate the SSI or TANF letter, capture the rent and utility figures. This is the use case where AI genuinely returns hours. A worker who spent fifteen minutes manually transcribing numbers from forty pages now reviews a structured extraction in three.
But extraction is also where the first failure can enter, and it is a quiet one. A model reading a blurry pay stub photo may transpose a digit, read a year-to-date figure as a period figure, or miss a second job entirely if the documents for it are at the back of the file. The opening case had a note that a second job ended in February. A model that read the most recent pay stubs and projected them forward as ongoing income would have overstated the household's income and could have pushed the family over a threshold they were actually under. The worker must treat the extraction as a draft to be checked against the source documents, the same claim-by-claim discipline used for any AI-drafted record, not as a finished data entry.
Stage Two: Policy Identification
The second stage is the one that decides cases, and it is the most error-prone. Here the workflow identifies which rules govern this household: which program, which eligibility pathway, which tests apply, which deductions are allowed, which exceptions are triggered. This is where the SSI child changes everything, where a state option overrides a federal default, and where a rule that changed last October may still live in a model's training data in its old form.
The model can help here by surfacing candidate pathways and flagging factors that may matter: "This household includes a member receiving SSI, which may trigger categorical eligibility; verify against current state policy." That is the model used correctly, as a flag and a prompt to the worker, not as an answer. The model used incorrectly is the model that states a definite policy conclusion, cites a regulation by number, and presents it with the same confidence whether it is right or wrong. A policy citation generated by a language model is a hypothesis to verify against the current policy manual, never a determination to copy.
Stage Three: The Calculation
With the data extracted and the governing rules identified, the third stage is the calculation: apply the deductions, run the relevant income tests, check the asset limit, compute the benefit amount. This is the stage closest to true arithmetic and the one where the model is most reliable, provided the inputs from stages one and two are correct. A calculation built on a misread pay stub or the wrong governing rule will be precisely wrong, which is worse than roughly wrong because precision reads as authority.
The verification move here is to confirm that the calculation used the verified data and the correct rule, and to sanity-check the result against what the worker's own experience expects. A determination that a household of four with one part-time income is ineligible should make an experienced worker pause and look again, because that result is unusual and unusual results are where errors hide.
Stage Four: The Human Determination
The fourth stage is the determination, and it is the line no model crosses. The decision to approve or deny benefits is a government action that triggers due-process rights: the household is entitled to a written notice, to the reasons for the decision, and to a fair hearing where they can challenge it. The worker, not the tool, makes this decision and owns it. "The system calculated ineligible" is not a determination; it is an input the worker considered before making one.
In practice this means the worker reads the AI-assembled package, confirms the extraction against the documents, confirms the governing policy against the current manual, confirms the calculation, and then exercises judgment. Judgment matters even here because eligibility cases contain ambiguity the rules do not fully resolve: a household-composition question, a self-employment income estimate, a disputed expense. The worker resolves those, makes the call, and signs it as their own professional determination.
Stage Five: Notice, Record, and Audit Trail
The fifth stage closes the loop. The determination becomes a written notice to the household that states the decision, the reasons, the policy basis, and the household's right to a fair hearing and the deadline to request one. The case record captures what was decided, on what basis, and that a human made the call. Where AI assisted, the record should reflect that an AI tool was used for extraction and calculation support and that the worker verified the output and made the determination. That disclosure is what makes the work defensible if the determination is later challenged at a hearing, where an advocate may ask exactly how the decision was reached.
A Worked Case, End to End
Return to the single mother of three from the opening. Walk the workflow stage by stage to see where it protects her and where a careless version would fail.
Intake and extraction. The model reads the forty pages and returns: household of four (mother plus three children), most recent biweekly pay stub showing $1,420 gross, a lease showing $1,100 rent, a utility shutoff notice, and an SSI award letter for one child. It also returns a projected monthly income that annualizes the recent pay stub. The worker checks the source documents and catches what the model glossed: the handwritten note and the February pay stubs show the second job ended, and the $1,420 figure already reflects only the remaining job. The model had not invented income, but it had not flagged the income change either. The worker corrects the projection to reflect current circumstances. Time spent: four minutes, against the fifteen it would have taken to transcribe everything by hand.
Policy identification. The model flags: "Household includes a child receiving SSI; categorical eligibility may apply; verify state policy." The worker opens the current state policy manual and confirms that in her state, the SSI-receiving member confers categorical eligibility, so the gross-income test does not apply. The model had flagged the right factor. The worker, not the model, confirmed the conclusion against the live source. Had she skipped this and let the model run the gross-income test, the family might have been wrongly denied if the recent pay had been higher.
Calculation. With categorical eligibility confirmed, the relevant calculation is the net-income test and the benefit amount, applying the standard deduction, the earned-income deduction, the shelter deduction tied to rent and utilities. The model computes a monthly benefit. The worker confirms the inputs match the verified data and the result is in the expected range for this household size and income.
Determination. Because the applicant marked the expedited box and the household has very low income and minimal resources, the worker confirms the expedited-processing criteria are met and the seven-day standard applies, not thirty. She approves expedited benefits and makes the determination her own. The total time on the case, with AI support and full verification, was roughly nine minutes, inside her eleven-minute budget and with the family correctly served.
Notice and record. The system generates a notice stating the approval, the benefit amount, the basis, and the household's hearing rights. The record notes AI-assisted extraction and calculation with worker verification and a human determination. If this case is ever audited or challenged, the trail shows exactly how a person reached the decision.
The Failure Modes This Workflow Prevents
It is worth naming, plainly, the specific harms a disciplined workflow is built to prevent, because each one has happened in the real history of automated benefits systems and each one is a due-process failure with a human cost.
The wrong denial from a misapplied rule. This is the categorical-eligibility failure: a family that qualified is denied because the workflow ran a test that did not govern their case. The prevention is stage two, policy identification verified against a current source. History offers a warning here. Automated benefits systems that applied rules without adequate human verification, including high-profile fraud-detection failures in public-assistance programs, generated waves of wrong determinations that took years and lawsuits to unwind. The discipline is not bureaucratic caution; it is the lesson those failures taught.
The wrong denial from a misread document. This is the extraction failure: a transposed digit or a stale pay stub pushes a household over a threshold. The prevention is stage one verification against source documents.
The silent automation of the decision. This is the most insidious failure because nothing visibly breaks. The worker, under a 200-case load and an eleven-minute budget, stops verifying and starts rubber-stamping the model's draft determination. The determination becomes the model's in fact even though a human name is on it. The prevention is structural: the workflow must reserve stage four as a real human judgment with enough time allocated to perform it, and supervision must watch for the drift from verifying to rubber-stamping. The time AI returns from extraction is what funds the verification; if that time is instead refilled with more cases, the safeguard collapses.
The undocumented decision. This is the failure that surfaces only at a fair hearing, when no one can reconstruct how a determination was reached or whether AI played a role. The prevention is stage five: a record and notice that make the basis and the human authorship explicit.
Building the Workflow Into the Real Unit
A workflow on paper is not a workflow in practice. Three things make the difference between a design that protects people and a design that becomes a rubber stamp under pressure.
First, the verification steps must be built into the tool and the process, not left to individual memory. A worker clearing a queue at speed will not reliably remember to check categorical eligibility on every case. The workflow should force the check: a required confirmation that the governing policy was verified against a current source before a determination can be entered, a structured extraction the worker confirms field by field, a notice that cannot generate without a recorded human determination. Discipline that depends on a tired worker remembering it is discipline that fails on the busy days, which are most days.
Second, time budgets must reflect verified work, not raw model speed. If a supervisor sees that the AI tool produces a draft in ninety seconds and resets the per-case time budget to two minutes, the math forces workers to skip verification. The honest budget accounts for the minutes verification takes. The legitimate payoff of AI here is real but bounded: it turns a fifteen-minute transcription-and-calculation task into a nine-minute verify-and-decide task, which is a meaningful gain and a sustainable one. It does not turn a case into ninety seconds, and any plan that assumes it does is a plan to deny people their due process.
Third, the workflow needs supervisory and equity oversight above the individual case. Supervisors should sample AI-assisted determinations to confirm verification is actually happening, and the unit should review aggregate patterns for disparity: are denials clustering in a way that suggests a systematic policy misapplication or a biased extraction failure affecting particular groups. A single wrong determination harms one family. A systematic error in the workflow harms everyone it touches, and only oversight above the case level can catch it.
This third condition is the one most easily skipped and the one whose absence does the widest harm, so it deserves a concrete picture. Imagine the extraction tool reads handwritten or non-standard pay documents less accurately than typed ones, and imagine that the applicants who submit handwritten documents are disproportionately low-wage, cash-economy, or limited-English households. No single worker reviewing a single case would ever see this; each case looks like an isolated extraction error. Only when someone looks at the pattern across hundreds of determinations does the disparity become visible: one group is being misread, and therefore wrongly denied, at a higher rate than another. That is exactly the kind of quiet, structural inequity the field has learned to fear, because it does not announce itself in any one file and it compounds with every case the tool touches. Equity oversight is not a compliance ritual bolted onto the workflow; it is the only instrument that can detect a harm distributed thinly across many families. A unit that audits individual cases for accuracy but never audits the aggregate for fairness has a blind spot precisely where the most damage accumulates.
Key Takeaways
- Eligibility work looks like arithmetic but is actually layered policy. The model is reliable at the mechanical parts (reading documents, calculating) and unreliable at the parts that decide cases (which rule governs, whether it is current), so the workflow must keep a human firmly on the governing-rule and determination steps.
- The AI-supported eligibility workflow has five stages: document extraction, policy identification, calculation, the human determination, and notice plus audit trail. Exactly one stage, the determination, is reserved for a person and can never be delegated to the tool.
- Categorical eligibility is the canonical policy trap: a household with an SSI or TANF member may be exempt from the gross-income test entirely, so a determination that runs the wrong test can be precisely calculated and still wrong. Policy identification must be verified against a current manual, not a model-generated citation.
- Extraction returns real hours (a fifteen-minute transcription task becomes a few minutes of review) but can quietly misread a blurry pay stub or miss an income change, so the worker verifies the extraction against the source documents claim by claim.
- The determination to approve or deny triggers due-process rights: written notice, the reasons, and a fair hearing. "The system calculated ineligible" is an input, never a determination; the worker makes and owns the decision.
- The most dangerous failure is silent: a worker under a 200-case load and an eleven-minute budget drifting from verifying to rubber-stamping the model's draft. The time AI returns from extraction must fund verification, not be refilled with more cases.
- History (automated benefits systems and fraud-detection failures that generated waves of wrong determinations) shows the cost of skipping human verification. The discipline is the lesson those failures taught, not bureaucratic caution.
- The workflow must be built into the tool and the unit: forced verification checkpoints, honest time budgets that account for verification, recorded human authorship, and supervisory plus equity oversight that catches systematic errors no single case review would reveal.
Skill.re