Measuring a Workplace Mental Health Program: Metrics That Mean Something

In this article
Open a typical workplace mental health dashboard and you will find counts: webinars held, employees trained, EAP flyers distributed, app downloads, wellness challenge sign-ups. These numbers are easy to collect, and they all go up when someone is working hard. What they rarely show is whether anything changed for employees.
Measuring a program properly does not require a research team. It requires being clear about what the program is supposed to change, choosing a small set of metrics that track that change, and reading them with some humility. What follows is a metrics table you can adapt, a way to pick from it, and answers to the questions HR teams ask most when they try to report on this work.
Start with the chain of effects
Every program rests on a chain of assumptions, whether or not anyone wrote it down. A manager training program assumes that training raises managers' confidence, that confident managers have more and better conversations, that those conversations lead people to seek help earlier, and that earlier help reduces time off and turnover.
Write the chain out for your own program, then put at least one metric at each link:
- Inputs: what you put in (budget, hours, coverage)
- Reach: who actually used it
- Experience: whether it felt useful and safe
- Conditions: whether the workplace changed
- Outcomes: whether people's working lives improved in ways you can observe
Most dashboards stop at reach. The more informative, harder-to-game evidence sits further down the chain. If you are unsure what your program really covers, our workplace mental health self-check is a quick way to see which of four areas are strongest and weakest before you decide what to measure.
The metrics table
Treat this as a menu, not a scorecard. No organization needs all of it.
| Metric | Layer | Typical source | Cadence | Read with caution because |
|---|---|---|---|---|
| Spend per employee on mental health support | Input | Finance, benefits | Annual | Spend says nothing about quality |
| Share of people managers trained | Input | Learning system | Quarterly | Completion is not competence |
| EAP counseling cases per 100 employees | Reach | EAP vendor, aggregate | Quarterly | Vendors define "use" differently |
| Behavioral health claims per 1,000 members | Reach | Health plan, aggregate | Annual | A rise can mean better access, not worse health |
| Days to first therapy appointment | Experience | Plan reporting, employee survey | Twice a year | Rarely reported unless you ask |
| Employees who know where to get help | Experience | Employee survey | Annual | Awareness without trust changes little |
| Comfort raising mental health with one's manager | Conditions | Employee survey | Annual | Sensitive to recent events |
| Survey items on workload and control | Conditions | Employee survey | Annual | Needs identical wording each year |
| Unplanned absence rate | Outcome | HRIS | Quarterly | Many causes beyond mental health |
| Short-term disability claims for mental health conditions | Outcome | Carrier, aggregate | Annual | Small numbers swing widely |
| First-year voluntary turnover | Outcome | HRIS | Quarterly | Labor market effects dominate |
| Return to work after mental health leave | Outcome | Leave administrator, aggregate | Annual | Very small counts in most firms |
A few rows deserve comment.
EAP use is the metric most often reported and most often misread. Vendors count different things: some count only counseling cases, others fold in website visits or calls to a legal or financial helpline. Ask your vendor for its definition in writing and for counseling cases broken out separately. When use looks low, the reasons usually involve awareness and trust more than need, a pattern covered in why employees don't use the EAP.
Days to first appointment is one of the most useful and least collected access measures. A generous benefit is worth little if the nearest in-network therapist has a six-week wait. Ask your health plan for behavioral health network adequacy data, and consider a survey item asking employees who sought care how long it took. Federal parity law bars group health plans from imposing stricter limits on mental health and substance use benefits than on comparable medical benefits, and the Department of Labor's parity guidance explains what plan sponsors are entitled to ask their carriers for.
Conditions measures are where programs succeed or fail. If employees still report unmanageable workloads and managers who change the subject, a strong EAP and a training catalog will only go so far. The WHO's guidelines on mental health at work give organizational interventions on working conditions a prominent place for exactly that reason.
Choosing six, not twenty
Six to eight metrics, spread across the layers, is plenty for most employers. One workable pattern:
- One input metric, such as manager training coverage.
- One or two reach metrics, such as EAP counseling cases and the behavioral health claims trend.
- One experience metric, such as days to first appointment or awareness of support.
- Two conditions metrics from your employee survey.
- One or two outcome metrics, such as first-year turnover or unplanned absence.
Tie the metrics to the actions you are actually taking. If this year's main investment is manager training, the comfort-with-manager survey item belongs on the list. If you have just widened the provider network or added virtual therapy, days to first appointment does.
Baselines, and good news that looks bad
Take a baseline before any major change, even a rough one. Without it you will be arguing from memory a year later.
Then expect some metrics to move in the "wrong" direction for the right reasons. As stigma falls and access improves, EAP use and behavioral health claims often rise. Survey disclosure may climb too, as people come to trust that their answers are safe. Leaders who are not warned in advance can read these shifts as failure.
Picture a hypothetical 900-person software company that launches manager training and a virtual therapy benefit in the same year. Twelve months later, EAP counseling cases have roughly doubled, behavioral health claims are up, and noticeably more employees say they would feel comfortable raising a mental health concern with their manager. First-year turnover is flat. A dashboard that treats claims purely as a cost line would flag a problem. Read along the chain, the picture is a program reaching people who previously went without help, with outcome effects that have not had time to appear. Writing that expected pattern down at the start, before the numbers arrive, protects the program from being judged on a single line.
Keeping program data private
Every metric in the table should be aggregate. Set a minimum group size for any breakdown, with five as a floor and ten as the safer choice, and avoid stacking filters in ways that could point to individuals. Employers should never receive identifiable EAP or claims data, and reputable vendors will not supply it. Under the ADA, any medical information an employer does hold must be kept confidential and stored apart from personnel files.
Reporting upward
Senior leaders do not need twelve charts. One page per quarter, laid out along the chain from input to outcome, with a sentence or two of interpretation per metric, is usually enough. Add what you plan to do next and which of the four key areas each action supports. When a number moves, say whether you think the change is real or noise, and why.
Frequently asked questions
Should we try to calculate a return on investment?
Be cautious. Most employers lack the sample size and data access to separate a program's financial effect from everything else going on. The widely cited WHO-led estimate of a fourfold return on investment in treating depression and anxiety came from a global, population-level model of scaled-up treatment, not from any single employer's program. A consistent trend across the metrics above usually persuades leadership more than a fragile ROI figure.
How often should we survey employees?
Once a year for a full set of conditions items is typical, with a short pulse in between if you are making significant changes. Surveying more often than you can act on the results wears down trust quickly.
What counts as a good EAP utilization rate?
There is no dependable benchmark, partly because vendors define utilization differently. Compare your own rate over time using one consistent definition, and check whether use is spread across sites, shifts and groups or concentrated in a few.
Can we look at health plan claims data?
Only in aggregate, de-identified form, usually through your benefits consultant or carrier. It is useful for trends in behavioral health use and cost. It is not appropriate for anything that could single out individuals, and in a small organization even aggregate cuts may need limits.
Our numbers got worse after launch. What now?
First check whether "worse" means more people are using help they always needed. Look at the conditions and outcome metrics before reaching a verdict. If conditions measures such as workload or comfort with managers are also slipping, the program may be missing the real problem, and an assessment of working conditions is the next step.
We have 80 employees. Is any of this realistic?
Yes, with adjustments. Small groups produce noisy numbers, so track a handful of survey items over several years and supplement them with structured conversations. Our resource library includes free survey instruments and guides that suit smaller employers.
Who should own the measurement work?
Ideally someone with a foot in both HR and data, often a people analytics lead or an HR business partner with a benefits background. What matters more than the title is that one named person owns the definitions, keeps them stable from year to year and writes the interpretation, rather than leaving each vendor to report its own numbers its own way.



