How do we test for unexpected user flows? Is there an AI…
How do we test for unexpected user flows? Is there an AI "chaos monkey" for UIs that can perform random but intelligent actions to try and crash the application?
Data as of Oct 5, 2026ChatGPT SearchGoogle AI ModeBased on 359 AI responses from ChatGPT Search and Google AI Mode·Reviewed by Dimitry Apollonsky ·
To perform random stress testing on your UI, use ChaosMonkey Tester for web-based browser automation or the Application Exerciser Monkey for Android mobile apps. While tools like Netflix Chaos Monkey simulate infrastructure failures, specialized UI stress testers focus on simulating unpredictable user inputs like clicks, scrolls, and typing to surface unexpected flows or crashes.
Yes, the concept exists and has evolved significantly. While Netflix’s classic Chaos Monkey targets backend infrastructure and cloud instances by randomly killing servers, the UI equivalent is a blend of traditional Monkey Testing and Autonomous AI Exploratory Agents.
Instead of just hammering random pixels until the app bursts into flames, modern AI-driven UI "monkeys" use vision and language models to act like curious, erratic, or adversarial users.
How Do We Test for Unexpected User Flows?
Testing for "unknown unknowns"—the chaotic, irrational ways actual humans interact with software—relies on a few distinct methodologies:
Traditional Monkey/Fuzz Testing (The Dumb Hammer)
What it is: Scripts that fire thousands of random clicks, taps, swipes, and keypresses into an app per minute.
For mobile, the native Android SDK Monkey Tool ( ) or iOS equivalent tools have done this for over a decade.
ChatGPT Search· Sep 26, 2026
Yes. There are a few approaches—some genuinely AI-driven, others closer to traditional UI fuzzing. Let me check the current tooling landscape and distinguish what actually works from tools that merely advertise “AI testing.”
The Catch: It doesn't know what a "save button" is versus a "delete account" button. It triggers infinite false-positive noise and rarely completes a coherent multi-step workflow.
Autonomous AI Exploratory Agents (The Smart Chaos Monkey)
What it is: Vision Language Models (VLMs) and LLM-driven agents that look at the UI dynamically, read labels, comprehend app states, and make deliberate (yet unpredictable) choices to probe edge cases.
How it works: Rather than following a rigid script, the AI acts like an "impatient user" or an "adversarial tester". It maps out a live state machine of your app, tries unexpected input lengths, double-clicks submission buttons while loading, skips required steps, or injects weird characters to see if the UI breaks or throws unhandled exceptions.
Is There an AI Chaos Monkey for UIs?
Yes. While still an emerging category moving fast into production engineering, several tools and approaches fulfill this exact role:
Visual/Agentic Web Probers & Research Frameworks: Emerging frameworks (like open-source academic and enterprise projects such as WebProber ) use Visual Language Models to navigate webpages by looking at raw pixels, clicking, scrolling, and hunting for dead links, UI freezes, or unhandled errors without a pre-written script.
mabl (Agentic Tester): Mabl incorporates intelligent auto-healing and autonomous discovery features designed to explore application paths and surface visual or functional anomalies without heavy engineering maintenance.
testRigor: Uses generative AI and Vision AI to interpret the screen from an end-user perspective, allowing tests to run dynamically across platforms while adapting to unexpected layout changes.
Custom LangGraph / Multi-Agent Swarms: Many engineering teams build custom internal PoCs using orchestration frameworks like LangGraph, deploying a "swarm" of specialized AI personas (e.g., The Rage Clicker, The Back-Button Abuser, The Slow Network User ) against a staging environment on every pull request.
If you're looking to implement this, tell me:
Is your app Web, Mobile (iOS/Android), or Desktop?
Are you trying to integrate this into a CI/CD pipeline (like GitHub Actions) or run it as a one-off exploratory audit?
I can recommend the most practical tool or setup architecture for your specific stack.
Yes. The closest thing today is agentic exploratory testing: an AI agent controls a real browser, observes the UI/DOM after each action, decides what to do next, and looks for violations rather than merely replaying a predefined script.
It’s not quite a mature, universal “Chaos Monkey for UIs” yet—but you can get surprisingly close.
What you actually want
A useful UI chaos agent would:
Explore
Click buttons, links, menus, tabs, dialogs.
Fill forms with unusual but valid/invalid data.
Navigate backward/forward and reload at inconvenient times.
Open/close things in unusual sequences.
Change viewport size.
Interrupt flows halfway through.
Reason about the current state
“I'm on checkout.”
“There is a modal blocking the page.”
“The user has entered an address but hasn't submitted it.”
“This button became disabled after the previous action.”
Generate adversarial actions
Double-submit.
Submit empty/partial forms.
Paste extremely long strings.
Enter Unicode/emoji.
Rapidly change selections.
Go back after a mutation.
Refresh during a multi-step transaction.
Open two competing dialogs.
Manipulate URL parameters.
Attempt actions in an unexpected order.
Define invariants rather than expected paths
This is the really important part. Instead of saying:
Click A → click B → expect C
tell the agent:
The application must never show an uncaught exception, lose saved data, expose another user's data, reach an impossible state, or leave the UI permanently unusable.
Minimize failures
If action sequence 47 causes a crash, the agent should try to reduce it to something like:
Login → open settings → change timezone → back → refresh → crash
That turns “AI did something weird” into a reproducible bug.
Existing tools
There are several pieces of this ecosystem already.
github.com is an open framework for web-agent research. It provides environments where agents can interact with browsers and includes benchmarks such as WebArena, VisualWebArena, WorkArena and MiniWoB. It's more of a research framework than a turnkey production chaos tester.
mabl.com has agentic browser testing. Its agent observes screenshots/DOM, plans an action, performs it, and reflects on whether the action succeeded. It is currently oriented more toward generating and maintaining purposeful tests than unconstrained chaos exploration.
testrigor.com explicitly supports AI-assisted exploratory testing, including a “tester tour” intended to try unexpected inputs and break things.
There are also emerging agentic-browser approaches using Playwright + browser-control agents; Google's current codelab, for example, demonstrates an agent controlling Playwright through BrowserMCP and natural-language instructions.
I'd build the "UI Chaos Monkey" slightly differently
If you're building this internally, I wouldn't let an LLM randomly click everything. That's entertaining but produces terrible signal.
I'd use a hybrid state-machine + AI explorer:
┌──────────────────┐
│ Current UI state │
└────────┬─────────┘
↓
┌──────────────────┐
│ Extract state │
│ DOM + screenshot │
│ URL + console │
└────────┬─────────┘
↓
┌──────────────────┐
│ AI action planner│
└────────┬─────────┘
↓
┌──────────────────┴──────────────────┐
↓ ↓
Normal action Chaos mutation
click/type/navigate double-click/back/reload/
invalid data/interruption/etc.
└──────────────────┬──────────────────┘
↓
┌──────────────────┐
│ Execute via │
│ Playwright │
└────────┬─────────┘
↓
┌──────────────────┐
│ Invariant checks │
│ crash? console? │
│ broken UI? data? │
└────────┬─────────┘
↓
failure?
↙ ↘
yes no
↓ ↓
minimize trace explore further
The AI should choose interesting actions, while deterministic code decides whether something actually constitutes a failure.
For example, give it:
Goal:
Explore this application for 10 minutes.
Allowed:
- Click/type/select/scroll
- Back/forward/reload
- Open/close dialogs
- Change viewport
- Submit forms
Prefer:
- Unusual action sequences
- Partially completed workflows
- Boundary values
- Repeated actions
- Interruptions
- Navigation during mutations
Never:
- Delete production data
- Send real emails
- Charge real credit cards
- Call production APIs
Report as a bug if:
- Unhandled exception
- Blank/frozen UI
- Broken navigation
- Data unexpectedly disappears
- Invalid state becomes reachable
- Accessibility trap
- Network request produces an unexpected 5xx
- Assertion/invariant fails
For every bug, minimize the action sequence and save:
- screenshot
- DOM
- URL
- console errors
- network errors
- video
- exact action sequence
The interesting bit: make the AI seek novelty
You can score UI states by novelty:
novelty =
new URL
+ new DOM structure
+ new visible components
+ new application state
+ unusual transition
Then have the agent preferentially explore states it hasn't seen.
That's much closer to coverage-guided fuzzing, except the input isn't bytes—it is user actions.
The agent deliberately searches for unusual transitions such as:
Onboarding Step 2
→ browser Back
→ reload
→ Dashboard
→ Settings
→ browser Back
→ Onboarding Step 2
Those sequences are exactly where conventional happy-path E2E suites tend to be weak.
One important distinction
I'd actually run three different modes:
Mode
Agent behavior
Regression
Deterministic known flows
Exploratory
AI chooses plausible but novel actions
Chaos
AI deliberately tries interruptions, invalid states and pathological sequences
That gives you reproducibility from the first mode and discovery from the latter two.
So, yes—the technology exists, but "AI Chaos Monkey for UIs" is still a useful product description rather than the name of one dominant tool. The particularly interesting engineering opportunity is combining an LLM/VLM browser agent with coverage-guided exploration, invariant checking, and automatic failure minimization.
Yes. What you’re describing is increasingly called agentic exploratory testing: instead of giving the test runner a fixed script, an AI agent observes the UI, chooses the next action, explores different states, and looks for failures.
There are now several ways to build this, but I’d separate “AI exploration” from traditional monkey testing.
What an AI UI chaos monkey should do
A useful agent would repeatedly:
Start from a clean or seeded application state.
Inspect the current page/accessibility tree.
Choose an action based on the current state:
click something
type unusual input
submit a form
go back/forward
open/close dialogs
resize/navigate
use keyboard navigation
rapidly repeat an action
intentionally violate expected workflow
Observe the result.
Decide whether the new state is interesting and continue exploring it.
Detect anomalies:
uncaught JS exceptions
failed network requests
4xx/5xx responses
blank/broken UI
invariant violations
unexpected navigation
stuck loading states
accessibility violations
crashes
Save the entire action trajectory, screenshot/DOM state, console/network logs, and ideally a Playwright trace.
Minimize the sequence to find the smallest reproducing action sequence.
That last part is particularly important. Finding “something broke after 37 random actions” isn't nearly as useful as:
Login → Settings → Enable X → Back → Delete → Cancel → Enable X → crash
The interesting part: you don't actually need a specialized product
Playwright + an LLM agent is probably the most flexible architecture today. Playwright now has official Test Agents: a planner can explore an app, a generator can turn plans into executable tests, and a healer can investigate and repair failing tests. Playwright also exposes an MCP server that gives agents structured browser control.
You can therefore build something like:
┌──────────────────┐
│ Your Web App │
└────────┬─────────┘
│
accessibility
+ DOM
│
┌───────▼───────┐
│ AI Explorer │
│ │
│ "What should │
│ I do next?" │
└───────┬───────┘
│
┌──────────┴──────────┐
│ │
Playwright App telemetry
│ │
▼ ▼
browser errors/logs
│ │
└──────────┬──────────┘
▼
┌───────────────┐
│ Failure Judge │
└───────┬───────┘
│
┌─────────▼─────────┐
│ Minimize + Replay │
└─────────┬─────────┘
│
▼
Reproducible test
The AI doesn't need to decide every individual click. A much better design is a hybrid:
Randomness gives you unexpected paths.
State awareness prevents the agent from wasting its entire budget clicking the same button.
LLM reasoning identifies interesting actions and unusual combinations.
Deterministic Playwright executes and records them.
Property/invariant checks determine whether something actually went wrong.
There are already projects doing pieces of this
For example, QA Agent is an open-source exploratory web-testing project that uses Playwright and optionally an LLM to explore applications through mouse, keyboard, forms, navigation, and accessibility testing.
AutonomousQA Agent similarly crawls an application, infers user flows with an LLM, generates Playwright tests, executes them, captures traces, and classifies failures.
There are also commercial products such as Wopee, which describe autonomous agents that explore an application, generate Playwright tests, and self-heal when the UI changes.
And there's active research specifically around making random UI exploration smarter. A 2026 study of LLM-assisted random GUI testing found that using the LLM to escape areas where random exploration gets stuck substantially improved coverage and crash discovery in its mobile-app experiments.
I'd make yours more "chaos monkey" than ordinary AI QA
The interesting implementation is something like a stateful fuzzing agent.
Give every UI state a fingerprint:
state =
URL
+ visible interactive elements
+ accessibility tree
+ important application state
The agent preferentially explores new edges and combinations, rather than merely generating random clicks.
For example, it might learn:
"Opening the modal normally works."
Then discover:
Open modal
→ press Escape
→ immediately click underlying button
→ reopen modal
→ browser Back
→ submit
That's precisely the kind of sequence conventional happy-path E2E tests tend not to contain.
Add "evil user" behaviors
I'd explicitly give the agent a mutation library:
Input mutations:
empty string
huge string
Unicode
emoji
whitespace
malformed email
negative numbers
0
MAX_INT
decimal where integer expected
Interaction mutations:
double click
rapid click
click while loading
submit twice
Back during request
reload during request
open two dialogs
keyboard-only navigation
Tab past boundaries
Escape at unusual times
Navigation mutations:
Back
Forward
reload
deep-link directly
open route in new tab
change URL parameters
Timing mutations:
action immediately after render
action during animation
action while network request pending
delayed action
repeated action
State mutations:
logout in middle of workflow
expire session
remove permissions
modify data in another tab
stale browser state
This gives you intelligent chaos, rather than simply:
That turns an AI "fuzzer" into something developers will actually use.
One caveat: current agents are not yet magical autonomous QA engineers. Recent research on agentic web testing still finds significant gaps in their ability to discover subtle real-world bugs reliably. arXiv So I'd use AI to expand the search space, while keeping deterministic assertions, telemetry, and replay as the authority.
If you're building this yourself, Playwright + accessibility snapshots + an LLM + a state/trajectory database + automatic sequence minimization is the architecture I'd start with.
Yes, you are describing an AI-driven UI fuzzing / autonomous exploration agent —essentially a Chaos Monkey for the frontend.
While traditional Chaos Monkey (by Netflix) kills backend microservices or database connections at random, a UI Chaos Monkey uses computer vision, Large Language Models (LLMs), or reinforcement learning to intelligently crawl your app, poke at weird UI states, input gibberish into forms, click random buttons out of order, and try to throw your frontend into an unhandled exception state.
How Testing for Unexpected User Flows Works Now
Testing for non-linear, chaotic user behavior historically relied on dumb "monkey testers" (like the Android Monkey or iOS Accessibility Inspector randomizers) that fired pure random coordinates at a screen. They usually got stuck instantly on a login page or an obscure modal.
Today, AI-powered autonomous testing tools act as "intelligent" monkeys. They understand what a button, dropdown, or text field is via DOM parsing or visual computer vision, meaning they can actually log in, fill out multi-step forms out of order, and explore deep state spaces.
Tools That Act as an AI "Chaos Monkey" for UIs
Mabl (Autonomous Testing / Auto-Healing):BrowserStack and platforms like Mabl use machine learning models to autonomously crawl web applications, discover new or unexpected pages/states, and report broken behaviors or visual anomalies without explicit scripting.
Applitools (Eyes & Autonomous Agents): Primarily focused on visual AI, but its underlying engine maps out UI states and detects when unexpected actions warp layouts or trigger uncaught rendering crashes.
Reinforcement Learning & GUI Fuzzers: In academic and high-security enterprise spaces, researchers use Deep Reinforcement Learning (DRL) agents specifically trained to maximize screen-coverage reward functions, systematically clicking and swiping through mobile and web apps to trigger memory leaks or app crashes.
Open-Source / Custom LLM Agents: Many engineering teams now spin up lightweight Playwright or Selenium scripts coupled with an LLM/VLM (Vision-Language Model) loop. The agent takes a screenshot of the UI, decides what the "most chaotic or unexpected" next click would be (e.g., clicking "Delete Account" immediately after hitting "Add to Cart"), and executes it.
Would you like to explore how to build a quick proof-of-concept LLM-driven chaos script using Playwright, or are you looking to evaluate a commercial autonomous testing platform for your stack?
Testing for unexpected user flows traditionally meant writing massive end-to-end (E2E) test matrices or deploying classic GUI fuzzers (like Android's Monkey tool for mobile apps) that blindly click random coordinates, type garbage text, and hope to trigger an unhandled exception or memory leak.
However, blind fuzzing is dumb: it gets stuck on login screens, loops endlessly on a single modal, or never discovers deep, contextual workflows (like adding an item to a cart and applying an expired coupon).
Today, the evolution of that concept is the LLM-powered autonomous GUI agent or intelligent UI fuzzer , which acts as a cognitive "chaos monkey" for frontends.
Is there an AI "Chaos Monkey" for UIs?
Yes. The industry has shifted from dumb random clickers to Vision-Language Model (VLM) and Large Language Model (LLM) agents that can "see" a UI, understand its semantic meaning, and intelligently decide what chaotic or unexpected action to take next.
Instead of randomly clicking pixels, these agents operate like a mischievous user who understands English, recognizes a "Delete Account" or "Checkout" button, and intentionally tries to break state management.
Top Tools & Approaches in This Space:
Autonomous Agent Platforms: Tools like AskUI use computer vision and AI models to interpret desktop, mobile, or web screens dynamically, allowing you to direct an agent to "explore and stress-test" interfaces without rigid script locators.
Generative AI QA Agents: Platforms like TestSprite AI leverage multi-agent frameworks to autonomously explore applications, map out state transitions, and unearth edge-case bugs.
Research & Open-Source Fuzzers: Academic and open-source projects leverage Reinforcement Learning (RL) combined with LLMs (often referred to in research as LLM-guided GUI exploration ) to dynamically generate event streams that maximize state coverage and force application crashes.
How to Test for Unexpected User Flows Today
If you want to implement an intelligent "chaos" strategy for your application's UI, combine these three modern paradigms:
AI-Driven Exploratory Testing (The Smart Monkey)
Point an autonomous agent tool at your staging environment.
Give it a high-level persona or prompt: "You are an impatient, confused user. Try to break the onboarding flow by clicking back mid-process, double-submitting forms rapidly, and navigating out of order."
Because the agent understands the UI elements via DOM-parsing or visual screenshots, it won’t get stuck on a standard login wall.
State Machine & Property-Based UI Fuzzing
Map your app into a coarse state machine (e.g., Logged Out → Cart → Checkout → Payment).
Use tools that inject random transition sequences—forcing jumps from Checkout back to Home , or firing concurrent API requests by clicking "Place Order" three times simultaneously.
Self-Healing Regression Guardrails
As your AI chaos monkey thrashes your app, traditional test scripts usually break because element IDs or classes mutate.
Pairing your testing suite with BrowserStack Self-Healing AI or similar tools ensures that when the UI adapts or shifts during exploratory chaos, the tracking layer dynamically heals locators so the test continues tracking actual application crashes rather than brittle selector errors.
If you'd like to narrow this down, tell me:
Is your application web, mobile, or desktop?
Are you looking for an open-source library to run locally or a managed commercial platform?
Minimize the reproduction — turn “37 random actions caused something weird” into “click X, delete Y, press Back, click X → crash.”
Persist the interesting flows as regression tests.
What exists today
There are several pieces of this already.
mabl.com is probably one of the closer commercial products. It can generate E2E tests from natural-language descriptions/user flows, recover from UI changes, analyze failures, and continuously maintain coverage.
applitools.com can crawl a site, generate tests, execute natural-language flows, and combine functional testing with visual validation. Their current material specifically discusses using AI for exploratory/negative testing rather than only happy paths.
github.com is interesting if you want to build this yourself. It's an open framework for web agents that can operate browsers and perform open-ended tasks. It is explicitly a research framework rather than a polished production testing product.
Google has also published a codelabs.developers.google.com, demonstrating the basic architecture: an LLM decides what to do and browser automation executes it.
But there's an important distinction
Most current “AI testing” products are closer to:
“Tell me a user journey and I'll execute/maintain it.”
What you're describing is more like:
“Here is my application. Explore it autonomously and try to break it.”
That's substantially more interesting.
I'd build the latter as an agentic state-space explorer rather than a random-click bot.
For example:
┌───────────────┐
│ Start app │
└───────┬───────┘
↓
Observe UI/state
↓
┌───────────────────────┐
│ Generate candidate │
│ actions │
│ │
│ click / type / back │
│ refresh / escape │
│ submit / cancel │
│ mutate URL / etc. │
└───────────┬───────────┘
↓
Choose "interesting"
next action
↓
Execute it
↓
Observe new state
↙ ↘
normal failure?
↓ ↓
continue capture trace
↓
minimize trace
↓
create regression
The interesting part is the action-selection policy.
Instead of:
randomly click something
you want:
“Given everything I've seen so far, what action or sequence is most likely to reach an unvisited state or expose an invariant violation?”
For example, the agent might notice:
“This checkout button became enabled after entering a credit card.”
It can then deliberately try:
double-click checkout
checkout → Back → checkout
modify quantity while checkout is loading
refresh during payment
open two checkout tabs
remove the item after reaching checkout
submit with expired session
press Escape during payment
navigate directly to the success URL
alter an ID in the URL
use browser Back after a POST
reload halfway through a wizard
That's chaos engineering guided by a semantic model of the UI.
I'd also give the agent explicit invariants
This makes it dramatically more useful.
For example:
Never:
- charge a user twice
- show another user's data
- allow checkout with an empty cart
- lose entered data unexpectedly
- leave an infinite spinner
- reach an authenticated page after logout
- produce an uncaught JS exception
Always:
- destructive actions require confirmation
- Back/Forward preserves valid state
- refresh doesn't corrupt the workflow
- errors have a recoverable UI
Then the agent isn't merely looking for crashes. It's looking for violations of application invariants.
That's much closer to a UI equivalent of a traditional chaos monkey.
A particularly powerful architecture
If I were building this today, I'd use:
Playwright → browser control
Accessibility tree + DOM + screenshots → state perception
LLM/VLM → action planning
Application telemetry → error detection
State hashing → exploration / deduplication
Property-based testing → input generation
LLM → semantic invariant generation
Trace minimizer → reproducible bug
Playwright test generator → regression test
And I'd put a risk budget around the agent:
Allowed:
✓ navigate
✓ click
✓ type
✓ refresh
✓ back/forward
✓ manipulate test data
✓ open/close dialogs
Restricted:
⚠ payments
⚠ deleting real data
⚠ sending email
⚠ external API mutations
Never:
✗ production destructive actions
The result could produce bug reports like:
Checkout duplicate-order bug
Discovered autonomously after 184 actions.
Reproduction:
Add item to cart
Checkout
Click "Place order"
Immediately click "Place order" again
Navigate Back
Return to checkout
Expected: one order
Actual: two orders
Console: no errors
API: two POST /orders requests
Minimal reproduction: steps 1–4
That is the version of an “AI chaos monkey” I'd be excited about.
And there's a real reason not to make it too random: purely random exploration has an enormous state space and spends most of its time doing useless things. The AI should be randomized strategically—maximize novelty, boundary conditions, state transitions, and invariant risk rather than randomness itself. Current research/tooling is moving in this direction, but the fully autonomous “give it a UI and let it hunt for bugs” experience is still less mature than conventional AI-assisted E2E testing.
If you're considering building one, I can sketch a concrete architecture for an MVP using Playwright + an LLM + a state graph, including the agent loop and how I'd make it minimize crashes into deterministic regression tests.
Testing for unexpected user flows and application-breaking chaos in user interfaces traditionally relies on Monkey Testing and UI Fuzzing.
Yes, the concept of an AI-powered "chaos monkey" for UIs—which blends the random destruction of traditional monkey testing with the smart, context-aware navigation of modern AI—is a rapidly growing reality in software testing.
How We Test for Unexpected User Flows
Traditional Monkey Testing ("Dumb" Monkeys): Tools fire completely random coordinates, clicks, swipes, and gibberish text into an application to see if it throws an unhandled exception or crashes. While great for stress-testing basic stability, they waste massive amounts of time clicking the same "Home" button or getting stuck in trivial loops.
Property-Based and State-Machine Testing: Engineers map out permissible states of an application and let a state-machine engine traverse permutations to find dead-ends or unhandled state combinations.
AI-Driven Crawlers & Vision-Based Agents: Modern tools use computer vision and Large Language Models (LLMs) to "look" at a UI, understand what a button does (e.g., recognizing a shopping cart vs. a delete account button), and intentionally try weird, exploratory paths.
Is There an AI "Chaos Monkey" for UIs?
Yes. They are generally referred to as AI-driven autonomous UI explorers or Intelligent Monkey Testers . Instead of blindly throwing events, they use visual models or LLMs to act like a chaotic, semi-rational user.
Prominent tools and approaches in this space include:
Appium / Android UI/Application Exerciser Monkey (with AI extensions): The classic Monkey Tool generates raw pseudo-random streams of events for Android. Modern test platforms layer AI over these inputs to cluster crashes and prioritize deep navigation paths rather than hitting dead ends.
Autonomous QA & Visual Crawlers (e.g., QApilot, Testim, Mabl, or open-source LLM browser agents): These platforms deploy agents that "crawl" web or mobile applications dynamically. They build a live knowledge graph of your app, figure out how to log in on their own, and intentionally try deviant paths (e.g., clicking "Back" during a payment checkout, double-submitting forms rapidly, or injecting unexpected unicode into fields) to isolate unhandled exceptions and UI freezes.
Would you like to explore building a lightweight custom AI chaos script using Playwright and an LLM, or are you looking for an off-the-shelf enterprise testing platform for web or mobile?
Yes. What you are describing exists, although the industry is still moving from random monkey testing toward AI-driven exploratory testing agents.
The basic idea is:
Instead of writing “click button A, fill form B, submit C,” give an agent a browser/app and a goal: “explore this product like a user and find broken states.”
These systems combine:
Monkey testing (random or semi-random input generation)
Exploratory testing (human tester intuition)
LLM agents (reasoning about what a user might try next)
Pure randomness produces too much noise. The better systems add:
Risk weighting
Spend more time on:
signup/login
payments
permissions
settings
data deletion
multi-step forms
State awareness
Remember:
"I am logged in"
"I have an empty cart"
"This modal is open"
"This user has no permissions"
Invariants
Define things that should never happen:
Examples:
"Checkout should never show success without an order ID"
"A deleted account should not still access private pages"
"A submit button should not silently do nothing"
Replayability
A bug report should say:
1. Open profile
2. Change email
3. Disable network
4. Click Save
5. App enters infinite loading state
Console:
TypeError: Cannot read property 'id' of undefined
Screenshot:
...
Where I would use one
A strong setup is:
Deterministic E2E tests for critical flows
(login, checkout, onboarding)
AI chaos testing nightly against staging
Traditional fuzzing for APIs and parsers
Production telemetry to feed discovered user journeys back into tests
AI monkeys are especially valuable because they find the "we never thought someone would do that" paths—the same class of bugs that often escape normal regression suites.
If you are building one internally, the most practical starting point today is usually Playwright + an LLM planner + a screenshot/error collector, rather than trying to create a fully autonomous tester from scratch.
Yes. The idea is real, but “AI chaos monkey for UIs” is still more of an emerging pattern than a single mature, dominant product.
The interesting version isn't truly random clicking. It's an AI-guided exploratory agent that deliberately generates weird, state-dependent user behavior and looks for violations of invariants.
What I'd build
Think of it as:
LLM + browser automation + state model + fuzzing + crash/error oracle
The agent gets a browser and a test environment and repeatedly:
Observes the current UI
screenshot
DOM/accessibility tree
URL/route
visible controls
console/network errors
Chooses an action
click something
type unusual input
hit Back/Forward
open/close modal
refresh
submit twice
navigate away during a request
resize the viewport
use keyboard instead of mouse
interact with something it "shouldn't" be able to interact with
Maintains a hypothesis about state
"The checkout is probably processing."
"This modal should be closed."
"The user is authenticated but their session may have expired."
"This button just caused a network request."
Looks for anomalies
uncaught exception
blank/white screen
infinite spinner
broken route
duplicate transaction
inconsistent UI state
unexpected navigation
accessibility violation
HTTP 500
console error
data disappearing
invariant violated
Minimizes the failure
This is particularly important. Instead of reporting:
"I clicked around randomly and the app crashed."
it should try to reduce the sequence to:
Login → Cart → Checkout → Back → Submit → Submit
and save the screenshot, DOM, console log, network trace, and video.
That's much closer to property-based testing + fuzzing + an autonomous browser agent than traditional UI test automation.
There are already pieces of this
ServiceNow's BrowserGym is probably one of the most interesting open-source foundations. It provides environments for browser agents and supports benchmarks such as WebArena, VisualWebArena, WorkArena and MiniWoB. It can drive browsers through actions such as clicking and filling forms and is explicitly designed for experimentation with web agents.
There are also commercial products such as Testim, which use AI for generating/maintaining UI tests and self-healing element locators. Its newer agentic testing capabilities can generate tests from natural-language descriptions. That's useful, although it's more AI test automation than a true chaos monkey.
And WebArena demonstrates that autonomous browser agents can already perform fairly complicated, multi-step interactions against realistic web applications.
The key distinction
I'd divide UI testing into three levels:
Approach
Example behavior
Scripted E2E
"Log in → add item → checkout"
Random monkey
"Click/type randomly for 500 actions"
AI chaos agent
"Explore the application, deliberately violate assumptions, and find states humans didn't anticipate"
The third is where things get really interesting.
For example, suppose your app has:
Cart → Checkout → Payment Processing → Success
A conventional E2E test verifies the happy path.
A chaos agent might discover:
Cart
↓
Checkout
↓
click Pay
↓
immediately click Back
↓
Forward
↓
click Pay again
↓
refresh
↓
click Pay again
And discover that the backend charged the customer twice.
That's vastly more valuable than random clicking.
I'd make the agent adversarial
Give it objectives such as:
"Find ways to make the application crash."
"Find a way to reach an impossible UI state."
"Try to submit the same operation twice."
"Try navigating away at every interesting point."
"Try stale sessions."
"Try invalid/partial input."
"Try using the UI faster than a normal human."
"Try browser Back/Forward in the middle of workflows."
"Try opening the same operation in two tabs."
"Try refreshing during every asynchronous operation."
"Try keyboard-only interaction."
"Try extreme viewport sizes."
"Try interrupting every modal/dialog."
"Try to violate authorization boundaries."
Then combine that with randomness.
For example, 70% of actions could be selected intelligently and 30% could be mutations/fuzzing. That gives you exploration without turning the test into meaningless button mashing.
The architecture I'd recommend
┌──────────────────┐
│ Your Web App │
└────────┬─────────┘
│
Playwright
│
┌────────▼─────────┐
│ Browser Agent │
│ │
│ screenshot │
│ DOM / a11y tree │
│ URL / state │
└────────┬─────────┘
│
┌──────────▼──────────┐
│ AI Action Planner │
│ │
│ "What should I try │
│ next?" │
└──────────┬──────────┘
│
┌─────────────┼─────────────┐
▼ ▼ ▼
normal mutation random
action action action
└─────────────┼─────────────┘
▼
execute action
│
▼
┌──────────────────────┐
│ Oracle Layer │
│ │
│ crashes? │
│ console error? │
│ HTTP 500? │
│ invariant broken? │
│ visual anomaly? │
└──────────┬───────────┘
│
▼
failure discovered
│
▼
minimize sequence
│
▼
reproducible test case
The oracle layer is arguably more important than the LLM. Without good oracles, the agent can wander around indefinitely and produce interesting-looking but useless reports.
That gives you something much better than "random UI testing": AI-directed state-space exploration.
And you can borrow techniques from traditional fuzzing: when a sequence finds a failure, mutate that sequence and explore around it.
The important caveat
Current browser agents aren't yet reliably good enough to just say:
"Here's my production website. Go crazy."
Recent evaluations still show substantial weaknesses in autonomous browsing, particularly around multi-step transactions and maintaining state across complicated workflows.
So I'd run this against an isolated staging environment with synthetic accounts/data, with strict limits on destructive operations.
But technically, yes—this is absolutely buildable today, and I think the most compelling implementation is not a conventional AI testing product. It's essentially:
If you're thinking about actually building this, I can sketch a minimal implementation architecture (e.g. Playwright + Python/TypeScript + an LLM) that autonomously explores a web app for 30 minutes and produces minimized reproducible bug reports.
testautomation.applitools.com uses AI for browser test creation/execution and combines functional testing with Visual AI. Its current emphasis is reliable, reproducible tests rather than random exploration.
Custom LLM-Powered Playwright/Selenium Agents: Many engineering teams now build lightweight internal "chaos agents" using an LLM coupled with a browser automation tool like Playwright. You prompt the LLM with: "You are a chaotic user who wants to break this app. Look at the accessibility tree or screenshot, pick a non-obvious action, and try to crash it." The loop runs iteratively, logging console errors and crashes.