<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Evaluation on Eigenform Articles</title><link>https://www.eigenform.ai/insights/tags/evaluation/</link><description>Recent content in Evaluation on Eigenform Articles</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Wed, 23 Sep 2026 00:00:00 +0800</lastBuildDate><atom:link href="https://www.eigenform.ai/insights/tags/evaluation/index.xml" rel="self" type="application/rss+xml"/><item><title>LLM-as-a-Judge: Complete Guide to Using LLMs as Evaluators</title><link>https://www.eigenform.ai/insights/llm-as-a-judge-complete-guide-to-using-llms-as-evaluators/</link><pubDate>Wed, 23 Sep 2026 00:00:00 +0800</pubDate><guid>https://www.eigenform.ai/insights/llm-as-a-judge-complete-guide-to-using-llms-as-evaluators/</guid><description>&lt;p&gt;&lt;strong&gt;TL;DR&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;When an AI grades another AI&amp;rsquo;s answers, it needs instructions on what a good answer looks like and how many points each part earns. LLM-as-a-judge is the name for this setup (one AI doing the grading instead of a person), and it&amp;rsquo;s the only realistic way to check thousands of answers instead of a few dozen.&lt;/li&gt;
&lt;li&gt;This only works if three things are true: the grading instructions actually catch the specific mistake they&amp;rsquo;re meant to catch, the AI grader isn&amp;rsquo;t quietly favouring one answer over another for reasons that have nothing to do with quality, and someone tested the grader on answers with a known correct score before trusting it on real ones.&lt;/li&gt;
&lt;li&gt;There are two ways to use an AI grader: score one answer on its own, or show it two answers and ask which is better. These answer different questions, and using the wrong one for the job is a common mistake.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong&gt;Key Takeaways&lt;/strong&gt;&lt;/p&gt;</description></item></channel></rss>