why does the same evaluation run produce different correctness scores
Summary:
We copied the HCM Data Security Assistant and created a 3 question evaluation based on correct results when run in debug. We submitted the evaluation twice, the correctness score by LLM in the first output is .55 with this LLM Feedback comment:
The response lists the correct roles but many security profiles are incomplete or differ from the reference, making it only partially correct.
And correctness score in the second is .95 with this LLM Feedback comment:
The response exactly matches the reference list of roles and security profiles, providing complete and accurate information.
What could cause the difference in results?
Content (please ensure you mask any confidential information):
Tagged:
0