PW Logo
Blog
Partners
Pricing
Client login

Performance Management

How to Choose Performance Management Software: 10 Questions Every HR Leader Should Ask Before Buying

August 17, 2026

Fourteen months after go-live, someone asks in a leadership meeting why nobody outside HR seems to be using the thing. 

The answers come fast, because everyone has one ready. Managers say it takes too long. HR says managers will not engage. Somebody mentions the HRIS integration never quite worked. Somebody else notes that most of the goals in there have not been touched since the first quarter. 

All of it is true. None of it was unknowable. Every one of those failures was sitting in plain sight during the demos, and every one would have surfaced if somebody had asked the right question. 

That is the trouble with how these evaluations run. They are organized around what a product can do, which is the one dimension where every serious vendor scores well. What separates performance management software is what it is like to live with, and a standard RFP barely touches that. 

The stakes also moved this year. The regulatory context is also becoming more demanding. The EU Pay Transparency Directive requires employers to make the criteria used to determine pay and pay progression accessible to workers, while AI used in certain employment-related decisions is subject to increasing scrutiny under emerging AI regulation. That raises the bar for performance records: they increasingly need to provide a credible evidence base for decisions about progression, reward and development.At the same time, AI used in employment decisions is attracting bias audit and transparency obligations in a widening list of jurisdictions. Both land in the same place: whether your performance record can actually justify the decisions made from it. A system that produces thin, unevidenced reviews used to be a quality problem. It is becoming a legal one. 

The ten questions: 

  1. What is actually broken, and what does fixed look like in two years? 
  2. What does a manager see on the day a review cycle opens? 
  3. Where does goal progress come from? 
  4. Where does skills data come from? 
  5. What does your AI decide, and what does it only suggest? 
  6. What proportion of feedback happens outside mandated cycles? 
  7. Can we change this ourselves, and can you configure our hardest case now? 
  8. What happens when someone changes manager partway through a cycle? 
  9. Who runs this internally, and how many hours a month does it take? 
  10. What happens to our data if we leave? 

Every question here is a version of one question: where does this system's information come from? 

Products that generate their own record behave differently from products that store whatever people type into them, and nearly every difference that shows up in year two traces back to that. For each one below: why it matters, and what separates an answer that is real from one that has been rehearsed. 

Start Here 

1. What is actually broken, and what does fixed look like in two years? 
Write this down before the first demo. It is the only thing that keeps an evaluation honest once vendors start showing you features you did not know you wanted. 

There are roughly three problems here and they need different solutions. Administrative burden, where the process works but costs too much. Hollow reviews, where the process completes and the content is thin. And no strategic line of sight, where leadership cannot connect any of it to business outcomes. Performance management software helps a lot with the first, somewhat with the second, and only indirectly with the third. If your actual problem is that managers avoid hard conversations, nothing on the market fixes it. You will automate a hollow process and get hollow output faster, on schedule, with excellent reporting. 

Then describe year two, not launch. Launch always looks fine. Everyone is paying attention, HR is chasing, the sponsor is visible, and completion rates are high in a way that means nothing at all. Describe year two in behavioral terms. Managers giving feedback without being asked. Goals that reflect what teams are actually doing. Development conversations that reference specific work. That paragraph is what you score vendors against, not the feature grid. 

Where the Information Comes From 

These four are the core. If you only ask four, ask these. 

2. What does a manager see on the day a review cycle opens? 

Make them show you the screen. Not a slide about the screen, not a curated demo account. The real thing. 

A weak answer is a blank form, maybe with last cycle's goals sitting alongside it and a text box waiting. That is a documentation tool. Reviews written into it get reconstructed from whatever the manager can remember, and memory skews hard toward whatever happened most recently and most loudly. 

A strong answer is an assembled picture. What this person worked on, what shipped, what feedback they received across the year, where their goals landed, how their skills moved. The manager's job becomes judgment instead of archaeology, which changes the conversation itself rather than just the paperwork around it. 

A weak system creates reconstructed performance: the manager has to remember, search, and write. A stronger system creates observed performance: the manager starts with an accumulated picture of work, goals, feedback and capability. 

There is a second reason this matters more than it used to. Under pay transparency rules now taking effect across Europe, employees can ask for the objective criteria behind their pay and progression, and employers have to show how those criteria were applied. A review built from recall is difficult to defend to an employee and harder still to defend to a regulator. It is also where the gap between the performance cycle and the pay decision becomes expensive rather than merely untidy. Ask any vendor what a defensible progression record looks like inside their system, and watch whether they have thought about it or are hearing the question for the first time. 

3. Where does goal progress come from? 

Ask who updates a goal, how often, and what happens when nobody does. 

If progress is a percentage someone types into a field, it will go stale. That is not a discipline problem and reminders will not fix it. Updating costs time the work is already demanding, so it slides, and by review season the recorded objectives describe a plan that stopped being relevant months ago. It is the quiet mechanism behind most OKR drift. Goals should not simply record whether work was completed. They should provide an ongoing signal about whether execution remains aligned with the outcome the business intended. 

Ask to see a real customer's goals dashboard with the last updated dates showing. That single request tells you more than an hour of feature discussion. 

4. Where does skills data come from? 

Competency frameworks are accurate the day they are finished and unreliable a quarter later, for the same reason goals are. 

A weak answer: an annual self assessment against a framework. You have bought a static matrix with better typography, and it will decay on exactly the schedule your spreadsheet did. 

A strong answer: capability inferred from what someone ships, resolves, leads and reviews, so the picture maintains itself. That is the whole difference between skills linked to execution and skills declared once a year, and it is the hardest thing in a demo to fake. You have to understand the difference between a skills inventory and skills intelligence. 

There is a request that settles it in about ninety seconds. Show me how one real person's skill profile changed over the last two quarters. A product without genuine skills data cannot show movement. It will show you the framework instead and hope you do not notice the swap. 

5. What does your AI decide, and what does it only suggest? 

Every vendor claims AI this year. The question is where it sits. 

AI on the output layer writes summaries, polishes phrasing, suggests wording. Mildly useful. It does not change what the system knows, so you get better written versions of the same thin information. AI on the input layer changes what gets captured in the first place, surfacing contributions, spotting patterns, connecting work to goals and skills. That distinction is the whole basis of how we have built AI into PossibleWorks, and it is worth putting to every vendor on your list. Ask which one they built. 

Then ask the governance question directly, because it has stopped being philosophical. AI involved in employment decisions now attracts bias audit, notice and transparency obligations across a growing set of jurisdictions, and performance evaluation sits squarely inside that scope. So: what does your AI decide as opposed to recommend? Can you produce documentation for an independent bias audit? Can you explain, for one specific employee, why the system surfaced what it surfaced? 

That last question is where a lot of products struggle, because explainability was not a design goal when the feature was built. Ratings, promotions and hard conversations should stay with humans, and a serious vendor will say so before you ask. While you are there, ask what stops signal capture from turning into surveillance. Your employees will ask, and you need an answer you can repeat without wincing. 

What It Is Like to Live With 

6. At your customers two years past implementation, what proportion of feedback happens outside mandated cycles? 

Every vendor has check in functionality. A lot of it amounts to a scheduled reminder with a form behind it, completed because the system insists rather than because anyone finds it useful. 

The number above is the honest one, and most vendors do not have it. Ask anyway. What they say when they cannot answer is informative. Some admit it, some redirect to completion rates, and the redirect is its own answer. Completion is the easiest thing in this category to measure and the least useful thing to measure. A stronger measure is whether meaningful performance conversations happen naturally because the system makes them easier and more relevant. 

Then run the friction test, because friction is usually why that number is low. During the demo, count the steps from logging in to giving one piece of feedback about one person. If it means finding the right module, picking a cycle and searching for a colleague, feedback will happen when the system demands it and at no other time. Performance processes die of small frictions far more often than they die of bad design, which is the unglamorous case for everything sitting on one screen. 

Ask what the system does unprompted, too. Does anything flag a goal that has not moved, someone who has not had feedback in months, a team where check ins quietly stopped? A product that only responds when you open it gets used as often as someone remembers to open it. 

7. Can we change this ourselves, and can you configure our hardest case right now? 

Two halves of one question, and the second half is where you learn something. 

Ask them to make a realistic change live, during the demo. Add a question to a review template. Change an approval flow. If it takes minutes, you have bought flexibility. If it needs a support ticket, a services engagement, or the phrase "we can scope that," you have bought a configuration you will be living inside for years. Neither is disqualifying. You just need to know which one you are signing. 

Then bring your genuinely awkward case. Matrix reporting where two managers both have a legitimate say. Multiple countries, different calendars, different legal requirements. A calibration sequence your leadership insists on. Contractors who need a lighter version of everything. Vendors are excellent at the standard path. The edge case is where you find out whether the product can accommodate the way your organization actually works, or whether your organization will have to adapt to the product.. 

8. What happens when someone changes manager partway through a cycle? 

A manager resigns in March. Her reports get split between two others, one of whom joined in February. In November all of them sit down for reviews with someone who was not there for most of the year being assessed. Every performance management system is built on an org chart, and org charts do not hold still. Ask for this on screen: a manager change partway through a cycle, then what the incoming manager can see of the work that happened before they arrived. Most products handle it badly. Either the review already underway gets stranded, or an administrator reassigns records one at a time. 

If your company restructures at all, this arrives inside the first year, and the answer decides whether you are holding a continuous record of someone's career or a pile of disconnected fragments. 

The Parts Nobody Asks About 

9. Who runs this internally, and how many hours a month does it take them? 

Buyers interrogate the manager experience relentlessly and almost never ask about the administrator's. Somebody will own cycle configuration, permissions, template updates and troubleshooting, permanently. In some products that is a few hours a month. In others it quietly becomes most of a person's job, usually discovered around the time the implementation consultant stops replying to emails. 

Ask what a comparable customer spends. Then ask a reference customer the same thing separately, because the two answers are not always the same. While you have the reference on the phone, ask what caused the delays in their rollout. Every vendor has a smooth timeline ready. What you want is the tail. 

10. What happens to our data if we leave? 

Rarely asked, occasionally expensive. Performance records carry real weight in disputes and dismissals, and increasingly in demonstrating that pay decisions were made on defensible grounds. Several years of that history is not something you want to discover you can only retrieve as a flat file with the context stripped out. Ask what format it exports in, whether historical reviews leave with their attachments and comment threads intact, and specifically what happens to any work signals captured from your other systems. That last part catches people out, because it is the newest kind of data in these products and the least likely to have an export path designed for it. 

Vendors confident about retention answer this without flinching. The ones who get uncomfortable are telling you something about how they expect this to end. 

How to Use These 

Do not send ten questions as a document. You will get ten written answers drafted by a marketing team, and written answers are useless here. The value is in watching someone try to demonstrate something live. 

Pick five or six for the demo, chosen for your situation, and insist on a screen for each. Save the operational ones for reference calls, where you will get straighter answers than the vendor will give you. And ask reference customers one question that is not on this list: what do you wish you had known before you signed? 

One more thing. Get three managers into two of the demos. Not the enthusiastic ones. The busy, mildly cynical ones with full delivery loads. They are the people who will decide whether this works, and their reaction in the first ten minutes predicts more than your scoring matrix will. 

Score everything against your year two paragraph from question one, not the feature grid. The grid is where every product looks identical, which is precisely what it is for. 

The Bottom Line 

Most performance management software evaluations get settled on a scoring matrix, and most scoring matrices measure the exact layer where these products are indistinguishable. Then eighteen months later somebody asks why nobody is using it, and the reasons turn out to have been visible the whole time. 

So stop asking what a product has. Start asking where its information comes from. A system fed entirely by manual entry produces thin data no matter how good the interface, because it can only ever know what busy people remembered to tell it. A system that builds its picture from the work itself starts somewhere else, and everything downstream inherits that difference. The reviews, the coaching, the promotions, the pay, and now the ability to defend any of it. 

That is the thinking behind PossibleWorks: capturing signals from the tools teams already use, connecting those signals to goals and skills, and bringing the resulting performance intelligence into the flow of work.. Make us answer these ten the way you would make anyone else. And whatever you end up choosing, hold it to the standard in question two. What a manager sees the morning a review opens will tell you most of what you need to know about what you have bought.