You're almost there! Please answer a few more questions for access to the Applications content. Complete registration
Interested in joining? Complete your registration by providing Areas of Interest here. Register

why does the same evaluation run produce different correctness scores

Summary:

We copied the HCM Data Security Assistant and created a 3 question evaluation based on correct results when run in debug. We submitted the evaluation twice, the correctness score by LLM in the first output is .55 with this LLM Feedback comment:

The response lists the correct roles but many security profiles are incomplete or differ from the reference, making it only partially correct.

And correctness score in the second is .95 with this LLM Feedback comment:

The response exactly matches the reference list of roles and security profiles, providing complete and accurate information.

What could cause the difference in results?

Content (please ensure you mask any confidential information):

Howdy, Stranger!

Log In

To view full details, sign in.

Register

Don't have an account? Click here to get started!