Agent Evaluations is G2's independent testing program for AI agents. G2 independently tests vendors' live agents on standardized, buyer-informed tasks. The results give buyers side-by-side evidence of agent performance and give vendors a way to demonstrate how their product performs under the same conditions as competitors.
Agent Evaluations is currently in beta. Evaluations are free to vendors during the beta period.
Understanding Agent Evaluations
For each AI agent category, G2 builds a simulated company and connects each vendor's live product to it. The simulated environment includes realistic customers, account and order data, a written support policy, and a set of tools the AI agent can use to take action.
G2 evaluates the complete AI agent experience, including the model, retrieval, actions, guardrails, workflows, and settings working together.
Each evaluation follows the same process:
- Configure the product using its knowledge base, actions, workflows, and documented settings.
- Run buyer-informed tasks, including realistic cases and meaningful edge cases.
- Capture a complete transcript and a log of the actions the AI agent took.
- Review the results with the vendor before publication. Vendors can flag setup mistakes for G2 to fix and re-run, but they can't change their score.
G2 refreshes evaluations several times a year, or sooner if a product changes materially.
For the full technical methodology, including category-specific test details, refer to G2's Agent Evaluations methodology.
Understanding evaluation criteria
G2 scores every evaluated AI agent against four criteria:
| Criterion | Judged by | What it measures |
|---|---|---|
| Accuracy | LLM-judged | Whether the AI agent's claims are supported by evidence |
| Policy Compliance | LLM-judged | Whether the AI agent followed the written policy it was given |
| Relevance | LLM-judged | Whether the AI agent stayed focused on the user's problem |
| Completeness | Deterministic | Whether the required outcomes actually happened in the underlying systems |
Each task in an evaluation carries a point value based on its complexity. An AI agent earns those points for a task it passes and earns zero for a task it fails. G2 sums these results into an agent's Overall Score.
Agent categories
G2 evaluates agents across the following categories:
- Customer Experience AI Agent
- AI SDR (Sales Development)
- Marketing AI Agent
- AI Legal
- Recruiting AI Agent
- Voice AI Agent
- Security AI Agent
- Accounting AI Agent
- Coding AI Agent
- IT Support AI Agent
Understanding evaluation results and G2 reviews
Agent Evaluations results are a distinct signal from G2 reviews and ratings. Buyers see three separate types of information when researching an AI agent:
- G2 evaluation results: Measured performance from Agent Evaluations testing
- G2 ratings and reviews: Buyer feedback submitted through G2 reviews
- Product facts and claims: Capabilities reported by the vendor
Agent Evaluations doesn't affect a product's G2 Score, Market Presence, or Grid® report placement.
Requesting an Agent Evaluation
If you're a vendor with an AI agent, you can request an evaluation using G2's Agent Evaluation request form.
After you submit a request, a member of G2's Agent Evaluations team will reach out to discuss next steps.
Evaluations are free to vendors during the beta period. G2 plans to introduce paid evaluations in the future.