Auto Evaluation
Auto Evaluations for Generation Datasets
For generation datasets, you have the option of using an LLM to score the evaluation.
This is called an auto evaluation, and if you choose this option, the application variant will be run against the evaluation dataset and then an LLM will annotate the results based on the rubric.
This will run as an asynchronous job in the background. When the evaluation is complete, you can view the updated status on the application page.

