Once a quarter, we sat down and rated team members on two dimensions: work performance and social behaviour. Each received a score from 1 to 10. The assessment was relatively informal, with room for individual judgement.

We used these ratings mainly to follow development over time. Looking across several quarters helped us notice who worked consistently and where our assessments were moving in a positive or negative direction. Preparing individual feedback conversations played a smaller role.

That overview was useful. The experience also showed us where the numbers needed more context.

A changing score needs a reason attached to it

The interesting question was usually what had prompted a change. Had someone become more reliable? Were we seeing a recurring difficulty, or reacting to one recent situation?

Looking back, I would preserve more of that reasoning alongside the ratings. A short sentence describing the observation would make the history easier to interpret months later. I would also flag changes in responsibilities or reviewers, because those can change the context of an assessment.

My practical takeaway: when a score changes meaningfully, write down what prompted it.

Two columns did not give us two independent assessments

We often struggled to separate performance from social behaviour. In practice, work performance tended to dominate the discussion and influence our overall impression of a person.

That is a limitation I would make explicit to anyone trying something similar. Giving a dimension its own column does not necessarily mean you are evaluating it separately.

I would now ask for a distinct explanation for each score: which results support the performance rating, and which observed actions in working with others support the second rating? If the same incident informs both, we should explain which aspect belongs to each.

My practical takeaway: ask for a separate explanation for each dimension.

The discussion needed more attention than the scoring

We usually had time to discuss only a few people in depth. In those cases, we could explore what had triggered a particular assessment. We rarely reached that level of detail for everyone.

With that constraint, I would select cases before the meeting: meaningful changes, different views between reviewers, and positive developments worth understanding. Collecting individual ratings beforehand would help make those differences visible.

For each selected case, I would start with one question:

What did you observe that the rest of us may not have seen?

My practical takeaway: record those observations before trying to settle on a shared number.