Agent Evaluations

Prev Next

Agent Evaluations is G2's independent testing program for AI agents. G2 independently tests vendors' live agents on standardized, buyer-informed tasks. The results give buyers side-by-side evidence of agent performance and give vendors a way to demonstrate how their product performs under the same conditions as competitors.

Agent Evaluations is currently in beta. Evaluations are free to vendors during the beta period.

Understanding Agent Evaluations

For each AI agent category, G2 builds a simulated company and connects each vendor's live product to it. The simulated environment includes realistic customers, account and order data, a written support policy, and a set of tools the AI agent can use to take action.

G2 evaluates the complete AI agent experience, including the model, retrieval, actions, guardrails, workflows, and settings working together.

Each evaluation follows the same process:

  1. Configure the product using its knowledge base, actions, workflows, and documented settings.
  2. Run buyer-informed tasks, including realistic cases and meaningful edge cases.
  3. Capture a complete transcript and a log of the actions the AI agent took.
  4. Review the results with the vendor before publication. Vendors can flag setup mistakes for G2 to fix and re-run, but they can't change their score.

G2 refreshes evaluations several times a year, or sooner if a product changes materially.

For the full technical methodology, including category-specific test details, refer to G2's Agent Evaluations methodology.

Understanding evaluation criteria

G2 scores every evaluated AI agent against four criteria:

Criterion Judged by What it measures
Accuracy LLM-judged Whether the AI agent's claims are supported by evidence
Policy Compliance LLM-judged Whether the AI agent followed the written policy it was given
Relevance LLM-judged Whether the AI agent stayed focused on the user's problem
Completeness Deterministic Whether the required outcomes actually happened in the underlying systems

Each task in an evaluation carries a point value based on its complexity. An AI agent earns those points for a task it passes and earns zero for a task it fails. G2 sums these results into an agent's Overall Score.

Agent categories

G2 evaluates agents across the following categories:

  • Customer Experience AI Agent
  • AI SDR (Sales Development)
  • Marketing AI Agent
  • AI Legal
  • Recruiting AI Agent
  • Voice AI Agent
  • Security AI Agent
  • Accounting AI Agent
  • Coding AI Agent
  • IT Support AI Agent

Understanding evaluation results and G2 reviews

Agent Evaluations results are a distinct signal from G2 reviews and ratings. Buyers see three separate types of information when researching an AI agent:

  • G2 evaluation results: Measured performance from Agent Evaluations testing
  • G2 ratings and reviews: Buyer feedback submitted through G2 reviews
  • Product facts and claims: Capabilities reported by the vendor

Agent Evaluations doesn't affect a product's G2 Score, Market Presence, or Grid® report placement.

Requesting an Agent Evaluation

If you're a vendor with an AI agent, you can request an evaluation using G2's Agent Evaluation request form.

After you submit a request, a member of G2's Agent Evaluations team will reach out to discuss next steps.

Evaluations are free to vendors during the beta period. G2 plans to introduce paid evaluations in the future.