Blog/

Automated Testing

What is Autonomous Testing? The 5 Levels Explained

Every vendor says "autonomous." This framework scores what a system actually does versus what a human still has to do, so a claim can be checked instead of taken on faith.

Sauce Labs Author

Sauce Labs Author

July 18, 2025

Ask five vendors what "autonomous testing" means and expect five different answers — a suggestion engine to one, a system that ships changes with nobody reviewing them first to another. That gap is why autonomy needs a framework instead of a marketing buzzword. This guide defines the term and maps it to five levels of autonomy based on what the human still does, including what each level requires to work.

What is autonomous testing?

Autonomous testing is software testing in which the system decides what to test, writes the tests, runs them, repairs them when the application changes, and reports what matters, with human intervention reserved for the decisions it cannot make on its own.

The four jobs an autonomous testing system takes on are deciding, generating, executing, and maintaining. In traditional test automation, each of those jobs belongs to a person. A test engineer decides what to cover, writes the test scripts, configures the pipeline to execute them, and repairs the scripts when the application changes. In autonomous software testing, the system handles some or all of those jobs, and the degree to which it handles them is what separates one autonomous testing platform from another.

What stays with people at every level: intent, acceptance criteria, and the judgement call on whether a change is a defect or a design decision. A testing system that generates test cases from a login form can check that the form accepts valid credentials, but it cannot decide that the login timeout should be 15 minutes instead of 30 because that is a product decision, not a testing decision.

The term arrived now because models can read an application and describe what it should do, rather than replaying what a person recorded. That capability is the foundation of what is AI testing in the broader sense, and autonomous testing is the direction it points toward. Agentic testing, covered in its own section below, describes how the system reaches a result. Autonomous testing describes how much the human is still involved.

How does autonomous testing differ from test automation?

Traditional test automation executes instructions a person wrote and does exactly that every time, while autonomous testing uses artificial intelligence to decide which instructions are needed and writes them itself, which means the test suite changes without anyone editing it.

Scripted automation works on repetition: A human writes the test, the machine repeats it, and the machine breaks when the application moves. The test scripts are stable until the application is not, and then a person fixes them. The value of traditional automation is that the test does exactly what was specified. The cost is that maintenance effort grows with every test added to the suite.

Autonomous testing works differently. The system derives test intent from the application and the change, and rewrites its own test scripts when the application changes. What that removes is most of the maintenance effort that consumes QA teams: the morning spent reading failure logs, identifying which selectors changed, and updating test scripts that were never real failures. However, this approach adds a reproducibility problem because a suite that regenerates is not identical run to run. A test generated on Monday may cover the same flow differently from a test generated on Tuesday because the model's context changed or the application state differed. 

The practical middle ground most teams end up in: autonomous generation, human review, deterministic execution. The system writes the test, a person approves it, and the CI pipeline runs the approved version the same way every time.

The five levels of autonomous testing

There is no ratified industry standard for autonomy levels in software testing. The model below describes autonomy by what the human still has to do, so a team can place itself without a vendor's help. Level 0 is manual testing without automation and is the baseline the levels are measured from.

LevelNameWhat the system doesWhat the human still doesWhat has to be true for it to workWhere most teams are in 2026
1Assisted authoringSuggests steps, selectors, or assertions during test creationDecides what to test, writes and owns every testAn existing framework and an AI assistant that can read the applicationMost teams describing themselves as doing AI testing are here
2Self-healing executionRepairs broken locators and updates test scripts after a UI changeReviews every heal, rejects incorrect ones, owns the pass/fail decisionAn audit trail, a confidence threshold, and a reviewerThe sweet spot for teams with a mature automation suite
3Autonomous test generationGenerates test cases from the application, requirements or code diffsReviews every generated test before it enters the suiteA review gate and enough domain context to catch what the model missedEarly adopters, on bounded flows
4Autonomous coverage and prioritizationDecides which tests matter for a given change, retires redundant tests, flags coverage gapsSets the risk policy, reviews coverage decisions on a cadenceProduction usage data and enough run history for the model to learn fromRare, and only on stable, well-instrumented applications
5Fully autonomous testingOwns the entire testing lifecycle, sets its own quality gates, escalates only what it cannot decideAudits the system on a cadence and handles escalationsTrust that no team has yet earned in production, plus audit-grade evidenceNo vendor currently offers full, unsupervised Level 5 autonomy, but Sauce Labs AURA provides the necessary release assurance for progressing toward higher autonomy with human oversight

Level 1: Assisted authoring

A person still decides what to test and owns every test, with the system suggesting steps, selectors, or assertions. This level requires nothing beyond an existing test automation framework setup and an AI assistant that can read the application. The real gain is authoring speed: A test that took an hour to write from scratch takes 20 minutes with suggestions. The maintenance burden is unchanged because the person still owns every repair.

Most teams that describe themselves as doing AI testing in their automation testing workflow exist at Level 1, which is not a criticism. Faster test creation matters, and it is the safest entry point because the human reviews everything before it runs.

Level 2: Self-healing execution

The system repairs broken locators and updates test scripts after a UI change, without being asked. A button moves from a class selector to a data-test-id, and the self-healing engine locates the new selector and updates the test. The testing process continues without a person editing the script.

Level 2 requires an audit trail where every heal is visible and reversible, a confidence threshold below which the system asks instead of acting, and a human who can reject a heal. The gain is no more mornings spent fixing red tests that were never real failures. The risk is a heal that keeps a test green against the wrong element, hiding a real defect behind a successful repair. Every automated repair needs to be visible in the test results so a reviewer can catch it.

Level 3: Autonomous test generation

The system generates test cases from the application, the requirements, a design file, or a code diff, and adds them to the test suite, shifting test creation from a human-authored activity to a machine-authored activity with human review.

The requirement is a review gate. Generated tests encode assumptions nobody stated: The model assumes a form field is required because it appeared in every training example, but the field is optional for a specific user role. Without review, generated tests produce false failures that erode trust in the test suite. The gain is coverage on flows nobody had time to automate. The limit is business rules: A testing system that has never been told what a valid discount is cannot test for an invalid one.

Level 4: Autonomous coverage and prioritization

The system decides which tests matter for a given change and retires redundant tests, flagging test coverage gaps against real user flows rather than against lines of code. This level of automated testing requires production usage data, a run history long enough for the machine learning model to learn from, and agreement on what risk means for this application.

The gain is pipeline time, because the full suite stops running on every commit and only the relevant tests run in the fast path. The risk is silent coverage loss: The model retires a test it considers redundant, and the scenario that test covered breaks in production. A periodic full regression testing run stays in the schedule regardless of how good the selection model becomes.

Level 5: Fully autonomous testing

The system owns the entire testing lifecycle, sets its own quality gates, and escalates only what it cannot decide. Level 5 requires trust that no team has yet earned in production, plus evidence good enough for a compliance audit.

Any vendor claiming Level 5 should be asked which of the four levels below it can demonstrate on a customer's application with real data. A realistic target for most organizations in 2026: solid Level 2 on the main suite, moving into Level 3 on a bounded set of flows with human oversight on every generated test.

How autonomous testing works

The autonomous testing levels describe what happens. This section describes how it happens so the claims can be evaluated rather than blindly accepted.

Machine learning and continuous learning

Models trained on run history, application structure, and past defects power the decisions at Levels 2 through 4. Continuous learning cycles update the model as the application changes: new selectors, new flows, new failure modes feed back into the training data.

The model drift that follows when the application changes faster than the model updates is the risk. A self-healing engine trained on six months of stable UI behavior will misjudge a major redesign because it looks like a hundred simultaneous defects. The trigger conditions that should force a retrain or a rollback: a spike in low-confidence heals, a drop in test accuracy, or a new release that touches more than a stated percentage of the UI.

Natural language processing for test scripts

The promise is requirements and user stories converted into runnable test scenarios. Natural language processing reads a sentence like "a user should be able to reset their password using their email address" and produces the steps: navigate to the reset page, enter the email, submit the form, check for the confirmation message.

Plain English test scripts that a non-engineer can read and challenge are the benefit. Where the translation breaks is implicit domain knowledge that was never written down: the password policy has a 90-day expiration that nobody mentioned in the story, or the email confirmation is rate-limited to one per minute. The natural language model writes what it was told, not what was assumed.

Self-healing and test maintenance

When a selector stops matching, the self-healing engine re-identifies the element using a combination of attributes: nearby text, position in the DOM, visual appearance, data attributes, and element type. The confidence score on the match determines whether the system heals automatically or escalates to a person.

The audit record for each automated repair should include the old selector, the new selector, the confidence score, the timestamp, and the test run that triggered it. Measuring the maintenance effort removed in engineer hours per sprint rather than in vendor percentages keeps the metric honest. A vendor that claims 90% reduction in test maintenance should be able to show the hours before and after, not a percentage derived from a sample application.

Change impact analysis and test selection

Mapping a code diff to the tests that cover the changed code is how test execution time drops without dropping test coverage. Risk-based ordering puts the tests most likely to fail first, so the feedback loop is as short as possible: the commit stage runs the high-risk tests in minutes, and the release gate runs the rest.

The full regression testing run that still happens on a schedule catches what the mapping missed, whether indirect dependencies that the change impact model did not trace, integrations that fail only when two changes land together, performance regressions that only surface under heavy load, etc. Test selection doesn’t replace full runs. It organizes and sequences them.

Sauce Labs AURA is the worked example for the mechanics described above: Sauce AI for Test Authoring generates durable tests built to survive UI changes from plain-language intent rather than code, so a requirement becomes a runnable test without anyone writing selectors. The platform executes tests (including visual and accessibility tests) on secure device clouds (Real Device Cloud and Virtual Device Cloud) before Sauce AI for Insights finds failure patterns and root causes from the same execution history the platform already has, drawing on more than 8.7 billion test executions. Best of all, AURA learns autonomously with human review and approval. 

What is agentic testing?

Agentic testing describes a software QA approach where an autonomous agent is given a goal and the ability to take actions, and it decides the sequence of steps itself rather than following a script. The difference from autonomous testing: Autonomy describes how much the human is still involved, whereas agentic describes how the system reaches a result. A Level 3 autonomous testing system may or may not be agentic. But an agentic testing tool must be autonomous because autonomy is a core defining characteristic of agentic AI systems. 

What an agent does in a test run

An agent reads the application state, decides the next action, takes it, and evaluates the outcome against the goal it was given. When the agent encounters an unexpected state, it recovers by choosing a different path rather than failing on the first unmatched selector. This is the behavioral difference from scripted automated testing: A script fails at the unexpected state, but an agent adapts.

The trace of decisions the agent produces is the audit artifact. Each step records what the agent saw, what it decided, what it did, and what happened. That trace is what a human reviews after the run, making the result explainable rather than a binary pass/fail from a system nobody can interrogate.

Where agentic testing works today

Exploratory testing passes over a new build, looking for the obvious break nobody wrote a test for. The agent navigates the application, exercises flows based on the UI it encounters, and reports what broke. Not a replacement for scripted regression testing, agentic testing catches the defect class that regression suites miss (e.g., the new flow that was never scripted or the edge case that nobody thought to cover).

Flows too variable to script reliably are the second fit. Search results, recommendations, personalised content, and other dynamic features change on every load. A scripted test fails because the content changed, not because the feature broke. An agentic test evaluates whether the feature works regardless of which content appeared.

Triage is the third: An agent that can reproduce a reported bug by navigating to the reported state, performing the reported action, and confirming the reported outcome saves the engineer the first 30 minutes of every defect investigation.

Where agentic testing costs more than it saves

Agentic testing runs are slower and more expensive than a scripted equivalent doing the same check because each step includes a model inference that a scripted test does not need. A scripted login test takes two seconds. However, an agentic login test can take 10 to 30 seconds, depending on the model and the application complexity.

Results vary between runs, so an agent is a poor choice for a release gate on a stable flow. If the checkout process hasn’t changed, a scripted test that checks the same five assertions every time is faster, cheaper, and more trustworthy than an agent that explores the checkout and may or may not cover those five assertions.

The evidence problem: An agentic run that passes for a different reason each time is hard to defend in a compliance audit. Regulated teams that need to show exactly which test covered which requirement on which release will struggle to use agentic testing as their primary testing method for gated flows.

What autonomous testing does not do

Autonomous testing does not remove human testers. Exploratory testing and usability judgement stay with people, and so does domain reasoning. A system that can generate test cases and repair broken test scripts cannot decide that a user flow feels confusing or that an error message is unhelpful — let alone that a new design communicates the wrong priority.

It does not know the business rules unless somebody encoded them. Autonomous test generation follows what the model learned from the application and the test data. If a rule was never documented, never tested, and never failed, the autonomous testing system will not know it exists until a customer reports the defect.

Autonomous testing does not fix a testing strategy that was already unclear. More generated tests against vague requirements produce more noise, not better test coverage. If the testing process has no ownership, no traceability from requirements to test scenarios, and no pass criteria, autonomous testing tools make the problem faster.

It also doesn’t eliminate test maintenance. It changes it, from repairing test scripts to reviewing repairs and auditing decisions. The repetitive tasks move, but the engineering judgement stays.

Finally, autonomous testing does not make a slow CI/CD pipeline fast on its own. Parallel execution and infrastructure still decide runtime. Autonomous test selection reduces what runs in the commit stage, but the full regression suite still needs enough execution capacity to finish within the feedback target.

Where autonomous testing fits in the testing life cycle

During development

Intent-driven tests generated from the code change, running on the pull request for quick feedback. Self-healing stays on, with every heal surfaced in the code review rather than applied silently. The gate rule: An unreviewed automated repair should never block or pass a merge on its own.

Functional testing at this stage covers the changed flow and its immediate dependencies. Performance testing does not belong in the pull request gate unless the change is to a known hot path because the overhead slows the feedback loop without adding signal.

During integration and regression

Test scenarios generated from code diffs, prioritized by change impact analysis, with the full regression testing suite running on the release gate rather than on every commit. Test coverage measured against real user flows rather than against lines of code. Load and performance testing runs scheduled inside the CI/CD pipeline, with regressions correlated against recent changes.

Continuous testing at this stage means the test suite adapts to each build rather than replaying a fixed set of test cases. Adaptation is where Levels 3 and 4 earn their value — and where the review cadence matters most.

After release

Production monitoring feeds the test selection model with what users do in production. Escaped defects are routed back into the suite as new test cases. The review cadence that keeps the machine learning model current as the application changes: monthly at minimum, weekly if the release cadence is daily.

The testing lifecycle does not end at the release gate. For an autonomous testing system, production behavior is the ground truth that training data is built from, and the loop from production back to the test suite separates an autonomous testing platform from automated testing with an AI add-on.

Autonomous testing tools and how to check a vendor claim

Every autonomous testing tool in the category claims some degree of autonomy. These questions separate claims from capabilities, and the full ranked comparison of AI testing tools covers individual vendors.

The questions to ask in a demo

Which of the five levels can the vendor demonstrate on your application rather than on a sample app? What happens when the system is not confident in a heal or a generated test, and who sees that? Do the generated tests export as readable code and run without the vendor, which is the portability test that matters if the contract ends?

The trial that settles it

Change a selector and a layout deliberately and record whether the system heals, fails, or passes when it should not have. Measure false failures per run in week one against week two since the first week always looks worse as the system learns the application. Count the engineer hours spent reviewing what the autonomous testing system did because that cost replaces maintenance effort rather than disappearing entirely.

Verifying self-healing claims specifically

Ask for the audit trail on every heal from a real customer run, with identifying detail removed. Ask what percentage of heals are rejected by reviewers, and treat a very low number as a warning rather than proof of accuracy. If reviewers reject almost nothing, they may not be reviewing. Ask how a heal is rolled back after it turns out to be wrong, and test the rollback during the trial.

Risk, governance, and trust

Evidence and auditability

The system must retain information about what was tested, the version of the system used, the test results, and this information should be kept for the duration required by compliance. Additionally, there should be a decision trace for agentic runs, allowing for explanations of pass/fail results months later. Auditors should be able to see not just the results, but also the path taken to reach those results. It's important to record human sign-offs against the release rather than the testing tool, as the tool serves as the method and the release is the final artifact.

When the system must escalate

Confidence thresholds determine whether a decision is routed to a person for review instead of being applied automatically. Certain types of changes always require human oversight, regardless of the confidence score, including changes to payment processes, alterations in authentication methods, modifications in data handling, and any adjustments in regulated areas. Additionally, there should be rollback plans in place for automated test changes that prove to be incorrect, allowing for a reversal with just a single command or click. 

Data privacy and model training

Controls on production data reaching a model, in either direction, are essential: where test data and application detail are stored, who can access them, and who can delete them. Additionally, it's important to address the residency question for regulated teams, which involves understanding where the data is stored and which jurisdiction's rules apply. Furthermore, we need to explore deployment options that address these concerns. 

Metrics that show autonomous testing is working

Maintenance hours per sprint, before and after, is the primary measure at Levels 1 and 2 and what tells a team whether the autonomous testing system is saving time or creating a different kind of work. Track it by the sprint, with a baseline from before the tool went live.

False failure rate per run, tracked weekly. This number should fall as the system settles into the application. If it doesn’t, the system isn’t continuously learning.

‍Escaped defects per release, which is the only number that says test coverage improved rather than just grew. If the autonomous testing system generates more test cases and the escape rate stays the same, the generated tests are not covering what matters.

Mean time to detect a regression and CI/CD pipeline minutes per commit. Review hours spent on generated tests and automated repairs, counted as a cost against the hours saved. Software quality improves when the net is positive: hours saved minus hours spent reviewing exceeds zero, and escaped defects fall.

The numbers to ignore: tests generated per day, total test cases in the suite, and any autonomy percentage a vendor supplies, which all measure activity, not outcomes.

Where Sauce Labs fits in autonomous testing

Sauce Labs AURA — AI-unified release assurance — is a closed-loop system that authors, runs, and analyzes tests with human oversight, built on over 8.7 billion test executions. "Full lifecycle" describes what AURA covers — authoring, execution, and analysis in one platform, rather than three disconnected tools — not what level of autonomy it operates at. 

The platform context that applies regardless of the current feature set: AI capabilities run against the same test executions that already produce Sauce Labs browser and real device test results, so the run history, the evidence, and the reporting sit in one place rather than in a separate autonomous testing tool.

Autonomous testing tools amplify what a team already has. They do not replace what a team hasn’t built.

Score your team's autonomy level and pick the next step

Score yourself against the five levels using one question each: Does a person write every test? (Level 0 or 1.) Does the system repair tests on its own? (Level 2.) Does it create them? (Level 3.) Does it choose which ones to run? (Level 4.) Does it decide the gate? (Level 5.)

If the answer is Level 0 or Level 1 today, the next step is coverage and a test automation frameworks decision. More testing methods and a stable suite come before autonomous anything.

At Level 1 moving to Level 2: Turn on self-healing for one test suite, require every fix to be reviewed, and measure rejected heals for a month. If the rejection rate is above 20%, the self-healing capability does not understand the application well enough. If it is below 5%, the reviewers may not be reviewing.

At Level 2 moving to Level 3: Pick two flows nobody automated, generate test cases for them, and review every generated test as if a new hire wrote it. Track the review hours as a cost, and compare them against the authoring hours the generated tests replaced.

At Level 3 moving to Level 4: Instrument production usage first because test prioritization without usage data is guesswork. The machine learning model needs to know what users do before it can decide what matters.

The four numbers to record before starting, whatever level the team is on: maintenance hours per sprint, false failures per run, escaped defects per release, and pipeline minutes per commit. Without that baseline, you can’t tell whether the autonomous testing system made things better or just made them different.

The rule for the next vendor conversation: Ask for a demonstration at the level above the one your team is on, on your own application, with your own test data. Any autonomous testing tool that cannot demonstrate on real data has not demonstrated at all.

On This Page

Free trial

Sign up Free

Start testing smarter with the world's largest continuous testing cloud — now with AI built in.

Start Testing

Keep reading.

VIEW ALL POSTS
Automated Testing
VIEW ALL POSTS