Image: slack.engineering · rights & removal
Agentic Testing: Where Agents Fit in the E2E Testing Stack
Reporting by Slack EngineeringRead the original at slack.engineering
Executive Summary
Facts Only
* Agent-driven E2E tests involved over 200 agentic workflows using Playwright MCP and CLI.
* Tests validate goals rather than fixed journeys: goal $\rightarrow$ agent adapts $\rightarrow$ verify result.
* Three execution models were tested: Agent + Playwright MCP, Agent + Playwright CLI, and Generated Playwright Tests.
* The Playwright MCP configuration showed a failure rate of 0% for thread reply and $\sim 12\%$ for search discovery.
* Generated Playwright Tests had an $\sim 8\%$ failure rate for thread reply but $\sim 48\%$ for search discovery.
* Average runtimes ranged from 3 minutes (Generated Tests) to 11 minutes (CLI).
* The cost analysis showed that execution model mattered more than the underlying LLM model, with MCP approaches generally costing less than CLI and Code Generation approaches for equivalent flows.
* Failures in CLI-based runs were often caused by authentication or navigation instability, not agent reasoning errors.
* Playwright MCP execution allowed easier parallelization compared to CLI-based runs.
Full Take
From the original · Slack Engineering
Abstract Agent-driven end-to-end (E2E) tests add a new exploratory layer to testing, but should they replace traditional deterministic tests? We ran more than 200 agentic E2E workflows using the Playwright MCP, Playwright CLI, and agent-generated Playwright tests in test workspaces using non-production data to find out how agentic testing could fit into both our and your testing stacks.Read the full story at slack.engineering
Sentinel — Human
The text presents a detailed analysis of agent-driven testing workflows, grounding its claims in specific experimental metrics regarding reliability, speed, and cost across different execution models.
