AI for HR Certification
Capable · M25 · lesson 25 of 28 · queued
Preview — browse every lesson free. Enroll to mark lessons complete, open partner links and save your progress. Login & enroll →
Recognizing Bad AI Output in People Contexts
📖
now learning

Recognizing Bad AI Output in People Contexts

15 min

Overview

An HR manager asks AI to draft a termination letter. It reads fine. Technically sound. Six months later, the terminated employee sues. The letter's wording created an unintended implication that the company was retaliating, when the actual reason was performance. The letter looked good. It was actually dangerous.

This is why recognizing bad AI output in HR is a survival skill. Bad output in HR doesn't just waste time. It creates legal exposure. It damages culture. It sends the wrong message to employees. It can backfire spectacularly.

This lesson teaches you to spot bad HR AI output before it leaves your desk. You'll learn the seven red flags that show up over and over. You'll see them side-by-side with good output. By the end, you'll have a fast mental checklist that catches problems before they become expensive.

Why This Matters for HR Professionals

HR documents live longer than you think. A policy you draft today becomes institutional memory. An email you send gets forwarded and discussed in Slack. A performance review gets attached to a personnel file and could be read by a lawyer. A recruiting email sits in someone's inbox and shapes their perception of your company.

Bad AI output in HR creates compounding problems:
- Legal risk: Inaccurate language about compliance, liability, or policy can create exposure.
- Cultural damage: Tone-deaf announcements or miscommunications damage trust.
- Hiring failure: Bad recruiting outputs hurt your brand with candidates.
- Employee confusion: Unclear policies create frustration and support tickets.
- Documentation risk: Vague or legally unsound performance reviews become liability if employment ends.

The stakes are higher in HR than in other functions. You're writing about people's lives: their jobs, their pay, their benefits, their development. That requires a higher standard.

Red Flag #1: Plausible but False Specificity

AI excels at sounding authoritative while being completely wrong.

Example: A recruiter asks AI to draft recruiting copy about FMLA eligibility.

Bad output: "We're proud to offer Family and Medical Leave Act (FMLA) benefits to all full-time employees. Eligible employees can take up to 12 weeks of unpaid, job-protected leave per year for qualifying events. Coverage includes personal illness, family illness, adoption, and certain military circumstances."

This sounds accurate. Specific. Professional. It's also wrong in several ways:
- FMLA is 12 weeks per 12-month period, not per year (different rules apply to calculating this)
- It's not "per year" in the way most people understand it
- Whether employees are covered depends on company size, tenure, and location
- This doesn't address state variations (some states have more generous requirements)

A candidate reads this, makes decisions about joining your company based on false information, and discovers the truth later. Now they're frustrated. Or they sue because the job posting made promises that don't hold up.

How to catch it:
- Anything about benefits, legal requirements, or policy should be verified against your actual documents.
- Ask yourself: "Does this match our handbook?" If you're not 100% sure, verify.
- If it's in the recruiting pipeline, ask: "Is every fact here something I'd defend in court?"

Tip: For any AI output about policies, benefits, or legal requirements, treat it as a draft that must be verified by your legal/compliance team before it goes anywhere near an employee or candidate.

Red Flag #2: Bias Markers and Discriminatory Language

AI picks up biases from training data. Not out of malice. Just because the data contained them.

Example: Job description AI output

Bad output: "We're seeking a young, energetic Marketing Manager who is excited to grow with our company. Must be comfortable in a fast-paced, demanding environment where stamina is essential. We're looking for a 'digital native' who understands social media intuitively..."

The problems:
- "Young, energetic" = Age discrimination
- "Digital native" = Age/generational discrimination
- "Stamina is essential" = Disability discrimination (implies ability requirements not actually necessary)

This isn't intentionally discriminatory. AI just absorbed patterns from training data and reproduced them. But it's still illegal. And it narrows your candidate pool to people with a certain profile.

Example 2: Performance review red flags

Bad output: "Sarah is an excellent team player who gets along well with everyone. She has a warm personality and is always cheerful in meetings. She shows real promise in her role."

The problems:
- This reads like you're evaluating her personality, not her performance
- "Warm," "cheerful," "team player" are coded language often applied differently to different genders
- Nothing specific about actual accomplishments or impact

How to catch it:
- Search for demographics: "young," "energetic," "digital native," "fit," "culture fit" (often code for "like us")
- Watch for gendered language: Words applied differently to men vs women (ambitious vs aggressive, emotional vs passionate)
- Spot personality language vs performance language: If it's about who they are rather than what they did, it's off

Important: Biased job descriptions and reviews aren't just unfair. They're illegal. Always run recruiting and review output through a bias lens.

Red Flag #3: Tone-Deaf Language for Sensitive Topics

AI doesn't understand context the way humans do. Ask it to draft a layoff email and it might produce something that sounds corporate and cold when the situation requires empathy.

Example: Layoff announcement

Bad output: "We are writing to inform you that your position has been eliminated effective immediately. This decision is part of a strategic restructuring initiative. Your final paycheck will include severance compensation according to our severance policy. Please contact HR for next steps."

This is technically clear. It's also cold and impersonal. Employees read this and hear: "You're disposable." Even if you're being generous with severance, this language damages morale for people staying.

Better output (human-written): "We're making a difficult decision: we're eliminating X positions, effective [date]. Your position is one of them. This isn't about your performance. It's about how our business is changing. We're offering [X weeks] severance plus [benefits]. You'll talk with HR this week about timeline and next steps. I know this is hard. I'm available if you want to talk."

Better because:
- It's honest about the why
- It separates "this person" from "this role"
- It shows empathy without being maudlin
- It's from a human voice, not a corporate entity

Example 2: Difficult feedback

Bad output: "During this performance review period, you have demonstrated several areas for improvement. Your punctuality and adherence to deadlines require immediate attention. Communication with team members needs enhancement. You should consider professional development in time management."

This is bureaucratic and vague. The employee reads it and either doesn't understand what to fix or gets defensive because it sounds like judgment instead of coaching.

Better approach: Show AI an example of how you deliver feedback when you want to be direct but constructive. Have the output include specific examples and actions.

How to catch it:
- Read it out loud. Does it sound like a human or a robot?
- Ask: "Would a leader I respect send this email?"
- For sensitive topics (layoffs, discipline, difficult feedback), bias toward more human, more empathetic, less corporate.

Red Flag #4: Missing Critical Context

AI produces output that's technically complete but missing information that's obvious to you but not obvious to the audience.

Example: Onboarding checklist

Bad output: "First week: Complete HR onboarding. Meet with manager. Review handbook. Set up systems access. Meet the team."

This is technically a checklist. But it's missing:
- Who owns each item? (Is HR doing system access or IT?)
- What's the timeline? (All day Monday? Over the week?)
- What if something isn't ready? (What if facilities isn't done with desk setup?)
- What are the actual systems? (Slack, GitHub, Google Workspace, whatever, be specific)

Good output includes: Specific tasks, owners, dates, contingencies, contact info if something goes wrong.

Example 2: Policy update email

Bad output: "We're updating our remote work policy. Please review the updated handbook and reach out with questions."

This tells people what changed, but not:
- Why did this change?
- How does it affect them specifically?
- When does it take effect?
- How do I request an exception?

Good output includes: Why (business reason or employee feedback), what changed specifically, timeline, how to adapt, how to request exceptions.

How to catch it:
- Pretend you're the audience. Could you follow this without asking questions?
- For procedural output: Who, what, when, where, why, how, is anything missing?
- For announcements: Would someone reading this understand why and how it affects them?

Red Flag #5: Fabricated Citations or False Authority

AI confidently cites things that don't exist.

Example: Benefits communication about retirement contributions

Bad output: "According to the Department of Labor, employees should contribute at least 15% of gross salary to retirement accounts for optimal retirement security. Our plan recommends this allocation..."

The problem: There's no DOL recommendation that says "15%." This is completely made up. But it sounds authoritative.

An employee reads this, makes financial decisions based on false information, and later discovers the guidance was wrong. They're frustrated. They might even have a legal claim if they lost money based on your company's false benefits guidance.

Example 2: Policy language that cites law

Bad output: "Under state law, companies must provide [X benefit]. We're excited to offer this to our employees."

The problem: You have no idea if this is actually required by state law. AI made it up. When an employee fact-checks it or an employment lawyer reviews it, it's wrong.

How to catch it:
- Any citation of law, regulation, or official guidance must be verified.
- If it says "according to [authority]," verify it's real and accurate.
- When in doubt, remove the citation. "We believe X is important" is safer than "The law requires X" if you're not certain.

Important: Never publish guidance about legal requirements (FMLA, ADA, state law, etc.) without verifying it against actual sources. False information creates liability.

Red Flag #6: Overgeneralization and False Universals

AI writes as if one approach works for everyone. In HR, that's rarely true.

Example: Onboarding template

Bad output: "All new hires should be paired with a buddy for the first month. Buddies should introduce new hires to the team, show them around, and answer questions."

The problem: This might work for office-based entry-level roles. It doesn't work for remote hires. It doesn't work for senior hires who need executive mentoring, not peer buddies. It doesn't work in distributed teams.

Good output acknowledges variation: "New hires in our [location] office are paired with a buddy. Remote hires get [different approach]. Senior hires get [different approach]. Buddies are responsible for X; managers are responsible for Y."

Example 2: Performance feedback language

Bad output: "Managers should provide feedback to all employees monthly. Feedback should follow the format: what they did well, what they could improve, and where they're heading. This approach works for all roles and levels."

The problem: Monthly feedback might be right for individual contributors. C-suite execs might prefer quarterly strategic conversations. High-performers might find monthly feedback patronizing. Struggling performers might need weekly check-ins.

How to catch it:
- Watch for words like "all," "always," "everyone," "should". They're often wrong in HR.
- Ask: "Is this true for every role/person/situation we have?"
- If not, the output needs qualification: "For [specific group], do X. For [other group], do Y."

Red Flag #7: Confidentiality Leaks and Privacy Violations

This is the one that can blow up in a hurry.

Example: Performance review feedback

Bad output: "Marcus struggled with his communication this year. He didn't update the team during the Q2 release, which created problems. The CEO mentioned this was an issue in leadership meetings. A few people also complained in 1-on-1s that he can be abrupt."

The problem: You've shared information that was confidential. Who complained? How does Marcus know? This creates drama.

Better: Frame it in terms of impact and behavior, not he-said-she-said. "During the Q2 release, the lack of communication created surprises for downstream teams. Marcus is aware this is an area to develop."

Example 2: Investigation summary

Bad output: "We investigated the complaint about discriminatory comments. Sarah reported hearing comments from the engineering team. Three people confirmed similar experiences. The investigation found that [specific comments people made]."

The problem: Sarah is now identified as the complainant (not confidential). The people quoted in the summary can be traced (not confidential). If this gets to an employment lawyer, it's now evidence that the company didn't protect confidentiality in an investigation.

How to catch it:
- Assume the document might be read by a lawyer or the person being investigated.
- Remove names, quotes, and specific identifying details.
- Use aggregation: "Multiple employees reported similar concerns" not "Sarah and three others said..."
- Any investigation output should be reviewed by legal before distribution.

Tip: Documents used in performance issues, investigations, or terminations might end up in court. Write them that way from the start.

Try This Now: Three Exercises

Exercise 1: Find the Red Flags

Read these three HR AI outputs. For each, identify: What red flag is present? What could go wrong if you used this?

Output A (job description): "We're looking for a dynamic sales professional with the energy and passion to drive results. If you're a self-starter who loves fast-paced environments and thrives under pressure, this is the role for you. We want someone hungry for success and willing to do whatever it takes."

Output B (policy update): "Effective immediately, all employees must work in office Mondays through Thursdays. Remote work Friday is a privilege, not a right. Employees must request approval at least one week in advance. Unapproved absences will be considered unexcused."

Output C (performance review): "Jennifer's performance this year has been exceptional. She's a joy to work with and brings incredible energy to the team. She's one of our best performers and really makes a difference."

(Answers: A, gendered language, overgeneralization. B, tone-deaf, no flexibility, no consideration for accommodations. C, personality-focused not performance-focused, vague, would require rewriting with specific examples and accomplishments.)

Exercise 2: Red Flag Scan Template

Create a simple checklist you can use when you review HR AI output:

  • [ ] Does this contain accurate information? (Verify any claims about policy, law, benefits.)
    - [ ] Are there bias markers? (Search for age, gender, ability language.)
    - [ ] Is the tone appropriate for the context? (Especially for sensitive topics.)
    - [ ] Is critical context missing? (Could someone follow this without asking questions?)
    - [ ] Are there false claims of authority? (Anything citing law or official guidance verified?)
    - [ ] Are there overgeneralizations? (Does this apply to everyone, or do we need variations?)
    - [ ] Are there confidentiality issues? (Could names be identified? Are quotes used inappropriately?)

Print this. Use it every time you review HR AI output.

Exercise 3: Rewrite the Bad

Take one of the "bad output" examples from this lesson. Rewrite it to fix the red flag. Show what you'd change and why.

Example: "We're seeking a young, energetic Marketing Manager..." → "We're seeking a Marketing Manager with expertise in [specific marketing skills]. You have [X years] of experience managing marketing campaigns..."

Why: Removed age markers, focused on actual qualifications, showed what we're actually hiring for.

Practical Application - "What to Do Monday Morning"


  • Print or bookmark the seven red flags:
    - Plausible but false specificity
    - Bias markers
    - Tone-deaf language
    - Missing critical context
    - Fabricated citations
    - Overgeneralizations
    - Confidentiality leaks

  • Create a 60-second review checklist for any HR AI output you're about to use.

  • Before any HR AI output leaves your desk, scan it against this checklist. If you see a red flag, don't send it. Ask AI to fix it.

  • Start a "bad examples" file. When you see problematic AI output (yours or someone else's), save it with notes on why it's bad. This becomes your reference library.

  • For sensitive documents (investigations, terminations, performance issues), always add a final review step: "Would a lawyer approve this?" If not sure, ask.

Key Takeaways

  • Verify specificity: Plausible-sounding information about policies or laws must be verified before use.
    - Run a bias check: Job descriptions and reviews are vulnerable to bias language, search for red flags.
    - Tone matters most in sensitive contexts: Layoffs, difficult feedback, policy changes need human empathy.
    - Assume information gaps: Think like your audience, what context is missing?
    - Never cite law or authority you can't verify: Remove false citations; use general language instead.
    - Watch for overgeneralizations: "All," "always," "everyone". These are usually wrong in HR.
    - Protect confidentiality: Assume documents could be read in court, write them that way.

FAQ

Q: Is all AI-generated content bad?
A: No. Some AI output is genuinely useful. But you need to check it for these red flags. Never assume it's fine just because it looks polished.

Q: What if I can't verify whether something is accurate?
A: Don't use it. Either verify it yourself, ask someone who knows, or remove that claim from the output. It's not worth the risk.

Q: How careful do I need to be with internal communications?
A: Treat everything as potentially document-able. Employees screenshot things, forward emails, save policies. Anything in writing might end up in a legal proceeding. Write accordingly.

Q: What should I do if I catch a red flag in AI output I've already sent?
A: Correct it quickly. Send a follow-up: "I want to clarify my earlier message about [X]. Here's the accurate information." The faster you correct, the better.

Q: Can I use AI for performance reviews and investigations?
A: For performance reviews, yes, but verify for the seven red flags. For investigations, be extremely careful. Investigation outputs are legal documents. Consider legal review before using.

What's Next

Now you know what bad looks like. Lesson 1.4 is about the decision framework: when do you trust AI output enough to use it as-is, when do you need to edit it heavily, and when do you need to throw it out and start over? It's about building your personal "AI trust matrix."