If you've spent any time researching mobile testing strategy, you've read the argument. Emulators and simulators are approximations. They can't tell you about battery drain, thermal throttling, or how your app behaves when a call comes in. They're slow. When something breaks, you can't tell whether it's your bug or the emulator's. Therefore, test on real devices.
Every one of those observations is, in isolation, true, and we'd make several of them ourselves. Our own documentation is explicit about what belongs on real hardware: memory consumption, CPU profiling, manufacturer-specific sensors, native ARM libraries, carrier network behavior, and pixel-perfect display rendering. We run a real device cloud of more than 10,000 Android and iOS devices precisely because those things matter.
The argument still has a structural problem. It evaluates emulators and simulators against a single criterion of fidelity to a physical handset and then declares them insufficient. That's like evaluating a unit test against the criterion "does it prove the system works end to end?" No. It was never supposed to. A unit test is cheap and deterministic, and it catches a specific class of defect earlier than anything else can.
Engineering leaders don't actually buy fidelity. You buy defects caught per dollar per unit of cycle time, subject to a risk ceiling you're willing to accept. Once you frame it that way, the question "are emulators good enough?" dissolves, and a much more useful one appears: which test, at which stage, on which substrate?
The critique is right about capability and wrong about cost
Here’s what the fidelity-only framing leaves out.
1. Most defects aren't silicon defects
Think about the last 20 production incidents your mobile team shipped. How many were caused by a chipset difference, a thermal characteristic, a hardware sensor, or a carrier network quirk? And how many were a null state on an empty list, a broken deep link, a regression in a payments flow, an auth token that didn't refresh, a layout that collapsed at a particular text scale?
In nearly every team we work with, the second category dominates, by a wide margin. Defects in that second group live in your code and your logic. They reproduce identically on a virtual device because the thing that's broken running on top of the OS, not the OS's relationship with the silicon underneath it.
Fidelity runs on a spectrum, and the correct amount is whatever your assertion actually depends on. Running a real device to confirm that a required-field error appears when someone submits a blank form buys precision that assertion never needed. The larger cost is capacity. A device-hour spent on string validation is a device-hour unavailable for the battery, sensor, and carrier network tests that can only run on hardware.
2. Under-fidelity and under-coverage: two risks, one budget
Here's the fragmentation picture as of August 2026, per Statcounter: no single Android version holds more than about 26% of worldwide share. Android 16 leads at roughly 26%, with Android 15 near 17%, Android 13 around 15%, Android 14 near 13%, and Android 12 and 11 adding another 18 percentage points between them. Your users are spread across those versions, across thousands of device models, and across OEM skins that each modify system behavior in their own way.
iOS consolidates far faster. Apple reported iOS 26 running on 79% of all active iPhones in June 2026 but "far faster" still leaves a long tail, and the tail is where the embarrassing bugs live.
Now do the capacity math. If every functional test has to queue for physical hardware, your matrix gets cut to whatever your device budget and session concurrency allow. Teams don't respond to that constraint by testing less thoroughly and admitting it. They respond by silently narrowing the matrix to the three or four devices they can afford to keep green, and then discovering the OS version-specific crash from a customer ticket.
A virtual device cloud loosens that constraint. Sauce Labs' Virtual Device Cloud supports Android 5.0 and up and iOS 14.0 and up, with no grid to maintain, which reaches further into the version tail than most real-device programs can economically cover. Our Real Device Cloud runs iOS 14.7.1 through iOS 27 and Android 10 through 17 on hardware built within the last handful of years, which includes many of the models your users actually hold. Breadth from virtual, depth from real, one platform, one matrix.
3. "Emulators are slow" is a claim with an expiration date
The performance critique dates from an era of x86 emulation of ARM instruction sets, where every guest instruction was translated on the fly and the overhead was brutal. ARM-native infrastructure removes that layer. The emulator runs ARM code on an ARM host, with no instruction translation and no x86-bridge ABI overhead.
Sauce Labs' premium ARM-native virtual infrastructure delivers 65% to 80% faster session starts and 49% to 60% faster test execution compared with earlier versions. ARM-based Android emulators are also available, covering Android 14 through 17. Session start time matters more than most leaders realize: It's the fixed cost paid by every test shard, so it's the number that determines whether a 400-test suite finishes in four minutes or 40. A merge gate that clears fast is one developers respect. A slow one is one they route around.
Four jobs virtual devices do that real devices structurally cannot
This is the part the fidelity argument can't account for because each one is a place where virtual devices carry the load well and where hardware time is better spent elsewhere.
Concurrency at merge gate scale. A physical device runs one session at a time and then needs to be reset. Virtual instances spin up on demand, in parallel, in the hundreds. If your goal is "every pull request runs the full smoke suite before merge," that goal is reachable on virtual and hard to fund on physical.
Determinism. The critique says simulators create ambiguity and you can't tell whether odd behavior is a real bug or an artifact. Fair. But that blade cuts both ways, and harder in the other direction. Real devices introduce more confounders, not fewer: A device at 12% battery throttles, a device with a stale OS-level permission dialog blocks the session, a device warm from the previous run behaves differently from a cold one, which is why Sauce Labs wipes every real device between sessions. Virtual devices start from a clean, identical state every single time. For regression testing, where the entire signal is "did this change break something that used to work?", that reproducibility is the point. Flaky tests waste compute, and worse, they teach engineers to ignore red builds, which is the most expensive failure mode in all of quality engineering.
Testing what ships next. When Apple ships a new iOS beta, the simulator is available on day one. The hardware running it in the wild isn't, and meaningful adoption is months away. Every team that has ever been blindsided by an OS release has been blindsided in exactly this window. Simulators and emulators are the only way to test against an OS that your users don't have yet, which means they're the only way to be ready before your users get it. Sauce Labs supports the newest iOS releases from day zero for exactly this reason.
Cost structure that lets you shift left. Maintaining an in-house grid and device lab runs into serious money. We've seen teams eliminate up to $200K a year in infrastructure maintenance by moving off self-managed infrastructure. But the deeper point isn't the line item. It's that when the marginal cost of a test run approaches zero, teams run tests earlier and more often, and a defect caught at commit time costs a fraction of the same defect caught in a release candidate. Cheap execution makes shift-left an actual practice rather than a slide.
The allocation model
Here's the frame we'd suggest replacing the debate with. Stop sorting by substrate. Sort by pipeline stage and risk class.
And a single rule your teams can apply without a meeting:
If the assertion depends on the physics of silicon, radio, battery, thermals, sensors, or a vendor's OS skin, then run it on a real device. If it depends on your code, run it virtual.
That rule is easy to teach and easy to audit. It routes roughly 80% of test volume to the cheap, fast, parallel substrate while ensuring that every hardware-dependent risk still touches hardware. In our experience, the teams with the best mobile quality outcomes have the clearest allocation policy, not the biggest device labs.
Which testing platforms offer both real devices and virtual clouds on one platform?
Sauce Labs runs a real device cloud and a virtual device cloud together on one platform. Perfecto and TestMu AI (formerly LambdaTest) also host real devices alongside emulators and simulators, while BrowserStack's mobile cloud centers on real devices.
A real device cloud hosts physical phones and tablets in data centers that you control remotely. A virtual device cloud hosts Android emulators and iOS simulators that start on demand from a clean state. Virtual capacity is what lets a pull request fan out across parallel sessions. Real hardware is what lets a release candidate meet actual batteries and carrier networks.
Teams whose entire suite depends on hardware behavior can be well served by a provider built around real devices. For everyone else, checking the box is the easy part. Four questions separate one platform from two products under one invoice:
- Does the same Appium, Espresso, or XCUITest suite run on both substrates with only a configuration change?
- Do results from real and virtual sessions land in one reporting view?
- Can virtual capacity scale to the concurrency your merge gate needs?
- Can you reserve dedicated real devices when a release gate calls for them?
Sauce Labs answers yes to all four: one suite, one reporting layer, parallel virtual capacity, and private real devices when a release gate needs them.
A note on who's telling you this
It's worth noticing the shape of the incentive. A platform whose primary business is real devices has a structural reason to argue that virtual devices are insufficient. A platform whose primary business is virtual devices has the mirror-image reason to argue they're plenty.
We sell both, a real device cloud and a virtual device cloud, at scale, on one platform, which leaves us no reason to talk you out of either. It also means the allocation model above is, for our customers, a configuration decision rather than a procurement decision: the same Appium, Espresso, or XCUITest suite, the same reporting, the same CI integration, with the target swapped per stage.
Does Sauce Labs offer both real devices and virtual devices on one platform?
Yes. Sauce Labs runs a Real Device Cloud with 10,000+ Android and iOS devices and a Virtual Device Cloud with Android emulators, iOS simulators, and desktop browsers. The same Appium, Espresso, or XCUITest suite can target either one, so moving a test between them is a configuration change.
Moving a suite from virtual to real shouldn't require a new vendor, a new contract, a new framework, or a rewrite. When it does, teams stop rebalancing, and a stale allocation is how you end up either overpaying for hardware you didn't need or shipping a thermal bug you should have caught.
Three moves for the next quarter
- Classify your existing suite against the physics rule. Most teams discover that a large share of what currently runs on real devices has no hardware dependency at all. Moving it buys faster cycles and lower spend, and it frees hardware time for the assertions that need it.
- Set a cycle time budget for the merge gate, not a device budget. Decide that the PR suite must finish in, say, under 10 minutes, then let that constraint drive substrate choice. It will drive you toward virtual, for the right reason.
- Protect a real-device slice in the nightly run and a real-device gate before release. Breadth on virtual is only safe if depth on hardware is non-negotiable. Write it into the release checklist so it survives a crunch.
Mobile testing is a portfolio decision
The mobile testing conversation has been stuck on a question of trust, “can you trust an emulator?” Engineering economics is the better question. You don't trust a unit test to tell you the system works. You don't trust a load test to tell you the UI is right. You compose them, because each one buys you a different kind of information at a different price.
Emulators, simulators, and real devices are the same kind of composition. Run them as a portfolio, allocate by risk, and stop paying real device rates to find out whether a button works.
Want to see what your current allocation is actually costing you? Request a demo or start a free trial and run the same suite across virtual devices and real hardware on one platform.






