Workplace Mental HealthAssessments & Screening

Measuring a Workplace Mental Health Program: Metrics That Mean Something

Measuring a Workplace Mental Health Program: Metrics That Mean Something
In this article
  1. Start with the chain of effects
  2. The metrics table
  3. Choosing six, not twenty
  4. Baselines, and good news that looks bad
  5. Keeping program data private
  6. Reporting upward
  7. Frequently asked questions

Open a typical workplace mental health dashboard and you will find counts: webinars held, employees trained, EAP flyers distributed, app downloads, wellness challenge sign-ups. These numbers are easy to collect, and they all go up when someone is working hard. What they rarely show is whether anything changed for employees.

Measuring a program properly does not require a research team. It requires being clear about what the program is supposed to change, choosing a small set of metrics that track that change, and reading them with some humility. What follows is a metrics table you can adapt, a way to pick from it, and answers to the questions HR teams ask most when they try to report on this work.

Start with the chain of effects

Every program rests on a chain of assumptions, whether or not anyone wrote it down. A manager training program assumes that training raises managers' confidence, that confident managers have more and better conversations, that those conversations lead people to seek help earlier, and that earlier help reduces time off and turnover.

Write the chain out for your own program, then put at least one metric at each link:

  • Inputs: what you put in (budget, hours, coverage)
  • Reach: who actually used it
  • Experience: whether it felt useful and safe
  • Conditions: whether the workplace changed
  • Outcomes: whether people's working lives improved in ways you can observe

Most dashboards stop at reach. The more informative, harder-to-game evidence sits further down the chain. If you are unsure what your program really covers, our workplace mental health self-check is a quick way to see which of four areas are strongest and weakest before you decide what to measure.

The metrics table

Treat this as a menu, not a scorecard. No organization needs all of it.

Metric Layer Typical source Cadence Read with caution because
Spend per employee on mental health support Input Finance, benefits Annual Spend says nothing about quality
Share of people managers trained Input Learning system Quarterly Completion is not competence
EAP counseling cases per 100 employees Reach EAP vendor, aggregate Quarterly Vendors define "use" differently
Behavioral health claims per 1,000 members Reach Health plan, aggregate Annual A rise can mean better access, not worse health
Days to first therapy appointment Experience Plan reporting, employee survey Twice a year Rarely reported unless you ask
Employees who know where to get help Experience Employee survey Annual Awareness without trust changes little
Comfort raising mental health with one's manager Conditions Employee survey Annual Sensitive to recent events
Survey items on workload and control Conditions Employee survey Annual Needs identical wording each year
Unplanned absence rate Outcome HRIS Quarterly Many causes beyond mental health
Short-term disability claims for mental health conditions Outcome Carrier, aggregate Annual Small numbers swing widely
First-year voluntary turnover Outcome HRIS Quarterly Labor market effects dominate
Return to work after mental health leave Outcome Leave administrator, aggregate Annual Very small counts in most firms

A few rows deserve comment.

EAP use is the metric most often reported and most often misread. Vendors count different things: some count only counseling cases, others fold in website visits or calls to a legal or financial helpline. Ask your vendor for its definition in writing and for counseling cases broken out separately. When use looks low, the reasons usually involve awareness and trust more than need, a pattern covered in why employees don't use the EAP.

Days to first appointment is one of the most useful and least collected access measures. A generous benefit is worth little if the nearest in-network therapist has a six-week wait. Ask your health plan for behavioral health network adequacy data, and consider a survey item asking employees who sought care how long it took. Federal parity law bars group health plans from imposing stricter limits on mental health and substance use benefits than on comparable medical benefits, and the Department of Labor's parity guidance explains what plan sponsors are entitled to ask their carriers for.

Conditions measures are where programs succeed or fail. If employees still report unmanageable workloads and managers who change the subject, a strong EAP and a training catalog will only go so far. The WHO's guidelines on mental health at work give organizational interventions on working conditions a prominent place for exactly that reason.

Choosing six, not twenty

Six to eight metrics, spread across the layers, is plenty for most employers. One workable pattern:

  1. One input metric, such as manager training coverage.
  2. One or two reach metrics, such as EAP counseling cases and the behavioral health claims trend.
  3. One experience metric, such as days to first appointment or awareness of support.
  4. Two conditions metrics from your employee survey.
  5. One or two outcome metrics, such as first-year turnover or unplanned absence.

Tie the metrics to the actions you are actually taking. If this year's main investment is manager training, the comfort-with-manager survey item belongs on the list. If you have just widened the provider network or added virtual therapy, days to first appointment does.

Baselines, and good news that looks bad

Take a baseline before any major change, even a rough one. Without it you will be arguing from memory a year later.

Then expect some metrics to move in the "wrong" direction for the right reasons. As stigma falls and access improves, EAP use and behavioral health claims often rise. Survey disclosure may climb too, as people come to trust that their answers are safe. Leaders who are not warned in advance can read these shifts as failure.

Picture a hypothetical 900-person software company that launches manager training and a virtual therapy benefit in the same year. Twelve months later, EAP counseling cases have roughly doubled, behavioral health claims are up, and noticeably more employees say they would feel comfortable raising a mental health concern with their manager. First-year turnover is flat. A dashboard that treats claims purely as a cost line would flag a problem. Read along the chain, the picture is a program reaching people who previously went without help, with outcome effects that have not had time to appear. Writing that expected pattern down at the start, before the numbers arrive, protects the program from being judged on a single line.

Keeping program data private

Every metric in the table should be aggregate. Set a minimum group size for any breakdown, with five as a floor and ten as the safer choice, and avoid stacking filters in ways that could point to individuals. Employers should never receive identifiable EAP or claims data, and reputable vendors will not supply it. Under the ADA, any medical information an employer does hold must be kept confidential and stored apart from personnel files.

Reporting upward

Senior leaders do not need twelve charts. One page per quarter, laid out along the chain from input to outcome, with a sentence or two of interpretation per metric, is usually enough. Add what you plan to do next and which of the four key areas each action supports. When a number moves, say whether you think the change is real or noise, and why.

Frequently asked questions

Should we try to calculate a return on investment?

Be cautious. Most employers lack the sample size and data access to separate a program's financial effect from everything else going on. The widely cited WHO-led estimate of a fourfold return on investment in treating depression and anxiety came from a global, population-level model of scaled-up treatment, not from any single employer's program. A consistent trend across the metrics above usually persuades leadership more than a fragile ROI figure.

How often should we survey employees?

Once a year for a full set of conditions items is typical, with a short pulse in between if you are making significant changes. Surveying more often than you can act on the results wears down trust quickly.

What counts as a good EAP utilization rate?

There is no dependable benchmark, partly because vendors define utilization differently. Compare your own rate over time using one consistent definition, and check whether use is spread across sites, shifts and groups or concentrated in a few.

Can we look at health plan claims data?

Only in aggregate, de-identified form, usually through your benefits consultant or carrier. It is useful for trends in behavioral health use and cost. It is not appropriate for anything that could single out individuals, and in a small organization even aggregate cuts may need limits.

Our numbers got worse after launch. What now?

First check whether "worse" means more people are using help they always needed. Look at the conditions and outcome metrics before reaching a verdict. If conditions measures such as workload or comfort with managers are also slipping, the program may be missing the real problem, and an assessment of working conditions is the next step.

We have 80 employees. Is any of this realistic?

Yes, with adjustments. Small groups produce noisy numbers, so track a handful of survey items over several years and supplement them with structured conversations. Our resource library includes free survey instruments and guides that suit smaller employers.

Who should own the measurement work?

Ideally someone with a foot in both HR and data, often a people analytics lead or an HR business partner with a benefits background. What matters more than the title is that one named person owns the definitions, keeps them stable from year to year and writes the interpretation, rather than leaving each vendor to report its own numbers its own way.