People who have worked with AI agents often have this confusion: the functions are all running smoothly, but no one can say for sure whether it is good or not. It's easy to get excited by just a few demonstration cases, but once they go online, users encounter a variety of failures. To answer this question, an evaluation system is needed.
The first step is to define tasks and success criteria. The output of an agent is often a process rather than an answer, so we cannot just look at whether the final response is correct or not. A feasible approach is to break down the task into several things: which tools to call, in what order, and whether the final product meets the constraints, and provide decidable criteria for each. The more specific the standards, the more reliable the evaluation.
The second step is to build an offline test set. Collect several typical tasks from real use, covering normal paths, boundary conditions, and known failure patterns. The test set doesn't have to be very large, but it needs to be stable - running the same set of test cases after each change can determine whether it is progress or regression. Be careful to avoid highly overlapping the test set with training or prompt examples, otherwise the evaluation will only repeat the demonstration.
The third step is to choose a scoring method. Write assertions that can be judged by programs, such as whether the tool is called correctly, whether the format is compliant, and whether the assertion passes, which is both cheap and objective; For parts that are difficult to program, such as whether the expression is appropriate and the reasoning is reasonable, model scoring or manual sampling review can be introduced. Combining two methods is more reliable than scoring alone.
The fourth step is to focus on process indicators, not just results. The number of tool calls, invalid retries, timeout rate, average steps, and single task cost are often indicators that expose problems earlier than pass rates. A task that barely passes but takes more than ten steps and burns a large number of tokens is essentially non extensible.
The fifth step is the online evaluation after going live. Offline test sets can never cover the long tail of the real world. After going online, signals can be continuously collected through user feedback, task abandonment rates, manual sampling, and other methods, and newly discovered failure cases can be returned to the testing set to form a closed loop.
It should be noted that the evaluation itself also has a cost. At the beginning, there is no need to pursue a big and comprehensive approach. Covering the most core links first and comparing whether changes can be made before and after is already a big step forward compared to iterating based on intuition.
Comprehensively organize industry information that has been publicly released.