What NGO Multiplayer Load Testing Actually Measures
NGO multiplayer load testing measures whether a game built with Unity Netcode for GameObjects remains responsive, synchronized, and operationally stable as more clients connect and interact. A useful test measures more than successful connections: it should expose server saturation, replication delays, bandwidth growth, allocation pressure, host authority bottlenecks, and the point where player experience begins to deteriorate. For most studios, the relevant capacity is not the absolute maximum number of sockets a server can accept, but the highest concurrency at which p95 response time, disconnects, and replication errors stay within explicit limits. Unity’s testing guidance supports creating multiple player instances in the Editor, while production testing also requires dedicated-server or container-based runs because Editor clients consume much more memory and CPU than packaged builds. A practical initial target might be 100 concurrent players in one match, 500 across five matches, or whatever the design actually requires. Those numbers are test hypotheses, not universal NGO limits. As of September 29, 2026, teams should validate their own build because NGO releases, transport behavior, hosting configuration, message frequency, and game-specific replication can change the result substantially.
Also worth reading: How Do Multiplayer Studio Operations Tools Reduce Launch and Live-Service Risk? · How Much Should a Studio Pay for Multiplayer Ops Platform Pricing Tiers in 2026? · What actually works for multiplayer server optimization in 2026, and how can a small studio improve performance without overspending?
The first step is to translate design expectations into pass or fail thresholds. For example, a session may allow 50 players per instance, require 20 combat participants, and need to sustain 10 instances during a launch event. In that case, a test at 200 total clients is meaningful even if no single instance contains more than 50 players. By contrast, a battle-royale test with 100 expected participants in one world requires a different replication and interest-management strategy. Teams should record connected-client count, server tick rate, client frame time, round-trip time, packet loss, bandwidth per client, spawn time, replication queue depth, and error count at one-second intervals. This produces evidence for scaling decisions rather than a vague claim that a server “handled 600 players.” Load, stress, and soak tests answer different questions: load testing verifies an expected workload, stress testing finds the failure boundary, and soak testing checks for leaks and degradation over time.
Build a Representative Test Instead of Connecting Empty Clients
A representative load test must generate traffic resembling the real game, not merely open the same scene hundreds of times. Idle clients are useful for connection and heartbeat tests, but they rarely reproduce the cost of movement, physics, combat, inventory updates, chat, projectiles, respawning, or late-join snapshots. The test harness should therefore run automated player agents or instrumented bot clients that execute a repeatable scenario: connect, join a match, move, rotate the camera, fire weapons, pick up items, take damage, die, reconnect, and leave at scheduled intervals. The scenario should reflect expected concurrency, session duration, and interaction density. If only 5% of players normally fight while 95% move, reporting an average based on 100% combat intensity can make capacity estimates misleading. Conversely, an event-only test that ignores navigation and global object replication can miss the most common form of server load.
Unity NGO includes multiplayer-player testing facilities for simulating multiple players in development, but those Editor instances are best used for functional and moderate-load checks. Production validation should use headless dedicated-server builds, cloud Linux instances, or containers, with real clients running on separate machines or appropriately partitioned compute. Each client should have a unique account or identity, use the same transport and protocol as production, and send actions on a controlled schedule. Deterministic input helps compare runs, while a small amount of randomized timing prevents the benchmark from benefiting from unnatural synchronization. Test data must also resemble production: large inventories, populated chat, active guild state, full match histories, and a world containing the maximum expected number of networked objects. A stripped development scene may fit more clients than the shipping game while telling the studio very little about its actual ceiling.
Instrumentation should be added before the first large run and kept lightweight enough not to distort results. Per-server metrics should include CPU utilization, resident memory, garbage-collection pauses, network egress, active connections, tick duration, command-processing time, and queued messages. Client-side telemetry should capture frame time, interpolation delay, time to spawn, ping, jitter, packet loss, and desynchronization indicators. A practical rule is to sample every second and retain raw time series, while logging detailed traces only for selected windows or failing sessions. If a test runs for 60 minutes and samples once per second, each metric series contains 3,600 points, which is enough to reveal early saturation, steady-state behavior, and gradual degradation. Store the build hash, NGO package version, Unity version, server configuration, region, instance type, scenario version, and test date with every result; otherwise, a later capacity decision can be based on an unrepeatable test.
Run the Test in Controlled Stages
The most reliable process begins with a smoke test at 5%, then 10%, 25%, 50%, 75%, and 100% of expected peak concurrency. Each stage should run long enough for connection, steady-state, and shutdown behavior to appear, although that duration will differ by game. A 10-minute stage may be adequate for an initial capacity scan, while a serious release candidate should also endure at least one hour at expected peak and preferably a multi-hour soak. Teams should preserve identical scenarios and thresholds between stages so capacity changes are visible. The server should begin with production-equivalent logging and observability, not maximum logging, because synchronously writing millions of individual replication events can itself become the bottleneck. Any temporary diagnostic setting must be documented and removed from the final benchmark.
During each stage, distinguish infrastructure limits from game-code limits. A server that reaches 95% CPU may need faster hardware, fewer replicated updates, less expensive physics, or a different match density. Constant memory growth over two hours suggests a retained object, collection, event subscription, or connection-lifecycle leak. Rising ping with stable CPU and bandwidth often points to network congestion, packet loss, region placement, or an overloaded relay. Increasing client frame time while server metrics remain healthy may indicate too much reconciliation, excessive visual work, or inadequate client interpolation. A spike at connection time often comes from expensive spawn snapshots or authentication fan-out rather than movement replication. The team should capture profiles during these transitions because a low average can hide a 2-second hitch that drives player churn.
A useful ramp profile combines sudden joins with gradual growth. Open 100 clients in 10 seconds to test lobby or matchmaking pressure, add 100 every minute to observe scaling, and then hold expected peak for 60 minutes. Include departures, timeouts, reconnections, host migration if supported, and client versions that are deliberately mismatched only when the production policy permits them. Run at least three representative repetitions when comparing instance sizes because shared-host noise can make one result appear superior. Median and p95 values are more informative than a single average: at a target of 500 clients, an average of 20 ms may coexist with repeated 250 ms spikes. A performance budget such as p95 server tick under 20 ms, p95 client frame time under 33 ms, fewer than 0.1% unexpected disconnects, and no monotonic memory growth can provide a defensible release gate, though the actual thresholds must match the game’s tick rate and genre.
Compare Infrastructure, Tools, and Self-Managed Testing
There is no universally best testing method. Open-source tools can provide control and low direct cost, managed services can reduce operational work, and custom harnesses are often necessary because NGO traffic is determined by game behavior. The supplied research identifies curl-loader as an HTTP, HTTPS, FTP, and FTPS loading tool, but it is not a substitute for an NGO gameplay client unless the game has a separate HTTP or WebSocket control surface. Unity’s Multiplayer Play Mode tooling is relevant for simulating players during development, yet it does not remove the need for realistic dedicated-server tests. AWS infrastructure can supply compute and network capacity, but the instance family, architecture, placement, storage, and operating model still have to be selected. The old reference to AWS m8azn instances should not be interpreted as a current NGO capacity recommendation; EC2 availability and pricing change, and the database or compute family best suited to one title may be outdated by September 2026.
| Feature | Unity Multiplayer Tools and Custom Bots | Cloud Load-Generator Service | Self-Managed Dedicated-Server Fleet |
|---|---|---|---|
| Traffic realism | High when using real game clients and actions | Medium, depending on supported protocol and orchestration | High with real clients and production builds |
| Direct cash cost | Often low for developer seats; custom engineering remains | Usually usage-based plus per-scenario or platform fees | Compute, storage, bandwidth, and staff time |
| Operational effort | Medium to high | Low to medium | High |
| Scalability | Limited by developer workstations or owned hosts | Convenient for temporary large populations | Flexible, but capacity must be provisioned |
| Measurement scope | Client, server, network, and gameplay | Platform metrics plus injected request metrics | Full observability and packet-level diagnosis |
| Best use | Functional simulation and repeatable gameplay scenarios | Connection, transport, and high-volume stress tests | Release certification and capacity planning |
Read NGO Capacity Signals Correctly
NGO does not provide one player-count number that applies to every game because cost depends on topology, bandwidth, replication, physics authority, and server implementation. A client-server design centralizes authoritative simulation, while distributed-authority or host-based configurations can distribute cost differently; NGO’s networking abstraction does not make all topology decisions automatic. The number of NetworkVariables, NetworkObjects, spawned entities, RPC frequency, tick rate, update rate, payload size, and interest-management scope can dominate capacity. A 100-player world in which every client observes every transform is fundamentally different from a 100-player world in which each client receives relevant entities within a bounded radius. The same hardware can support either pattern, so a claim that “NGO supports 1,000 players” is incomplete without workload and configuration details.
Connection acceptance is a particularly weak success metric. A server can accept 1,000 sockets while taking several seconds to spawn players, missing replication deadlines, or sending outdated states. Better gates include p50, p95, and p99 connection-to-playable time; the proportion of clients that enter the intended initial state; server tick overrun rate; outbound bandwidth per active client; and the number of correction or timeout events. For a 60 Hz client presentation target, the team might track frame-time consistency, but the server tick should be selected independently based on simulation requirements rather than copied automatically from render frame rate. NGO versions and transport implementations continue to evolve, so package documentation and the exact project configuration should be consulted for current APIs and known constraints. Benchmarking the release candidate remains the only credible way to state a supported concurrency figure.
Security and fairness also belong in a load program. Synthetic accounts should not bypass authentication, matchmaking, entitlement checks, or anti-cheat systems if those services are part of the production path, because skipping them creates unrealistic results. Load tests should use approved non-production identities, isolated queues, rate limits, and observability labels, and they must never interfere with live players. Randomizing only packet arrival does not make a test safe if scripts can issue real purchases, send messages, or mutate shared persistence. Schedule high-volume tests outside public launch windows, cap total traffic, and terminate it automatically when error rates, CPU, memory, or cost exceed defined limits. This reduces the chance that a capacity experiment becomes an outage, particularly when the tool can create 5,000 clients in minutes.
Avoid the Mistakes That Produce False Capacity Numbers
The most common error is testing the wrong binary. Editor performance, development symbols, verbose logging, local networking, and unoptimized scenes can make results substantially worse than a release build, while a stripped benchmark can make results falsely optimistic. The test must use the same serialization, network code, scene assets, tick settings, quality levels, and persistence path as the candidate release, with only observability carefully controlled. Another error is scaling all clients from one physical machine; CPU scheduling, memory bandwidth, loopback networking, and NIC limits can contaminate both clients and servers. Place generators in separate processes, hosts, or regions as required, and document whether client and server hardware are isolated. Running the server and load generator together is acceptable for a smoke test, not for final capacity certification.
Teams also make mistakes by averaging away failures, changing the scenario midway, and declaring success from connection logs. A benchmark that falls from 40 to 10 players per second is no longer a valid 1,000-concurrent-player test, even if every client eventually connected. Use stable arrival rates and record gaps. Do not count repeatedly reconnected clients as unique concurrent users without disclosure, and do not let reconnect logic mask an unstable first connection. Warm-up periods should exclude cold JIT, asset loading, cache population, and matchmaking delays from steady-state results, but those phases should be measured separately if they affect player experience. Comparisons should use identical client distributions and time windows.
A third mistake is ignoring the upper bounds imposed by hosting platforms and relays. A dedicated server may fit its simulation budget on one VM while regional transport, NAT behavior, socket limits, or provider quotas prevent that many real clients from reaching it. Conversely, passing a cloud relay test does not prove the authoritative server can sustain the resulting load. Test the full path in the intended region, including DNS, certificate handling, gateway or relay capacity, observability, and persistence dependencies. If the game uses a relay rather than direct client-to-server connectivity, report relay and server capacity separately. These distinctions matter to B2B game teams because a launch plan based on only one layer can fail under plausible conditions.
Decide When to Scale, Optimize, or Change the Design
Scale infrastructure when measured demand approaches a stable limit but the game code has adequate performance headroom. For example, if CPU remains below 70%, memory is stable, p95 tick time is under 20 ms, and adding clients mainly increases linear bandwidth, moving to larger or additional server instances may be the lowest-risk response. Use at least roughly 30% headroom for traffic bursts, node loss, deployment variance, and noisy-neighbor effects; that percentage is an operational cushion, not a law. Autoscaling should be based on live demand and game-relevant signals, not CPU alone, because an authoritative server can remain at 60% CPU while its tick budget is already violated. Measure the minutes required to add capacity and whether new instances can join the same lobby, regional pool, or party system without disrupting existing players.
Optimize before scaling when cost grows faster than player count, a hot path is expensive, or optimization has predictable gameplay impact. Common opportunities include reducing RPC frequency, batching reliable updates, using unreliable events for replaceable state, tightening NetworkVariable write conditions, culling irrelevant NetworkObjects, limiting ownership transitions, pooling spawned objects, and avoiding per-frame allocations. Field-of-interest management or partitioned worlds can reduce replication, but they introduce visibility, authority, and consistency concerns that require functional testing. Moving from 60 to 30 server ticks may reduce CPU and network cost, yet it can make movement or combat feel less responsive; it is a design decision, not a free capacity fix. Compare two builds at the same client count and scenario before claiming an improvement, and report latency as well as throughput.
Change the architecture when expected concurrency exceeds what centralized servers can provide economically or technically. At that point, consider sharding matches, separating simulation from social services, partitioning persistent worlds, adopting a distributed authoritative model, or limiting synchronized entity scope. This is a larger commitment and can introduce cheating, synchronization, persistence, and debugging problems, so it should follow evidence rather than a marketing target. Teams should act before launch when the expected peak is more than about 70% of proven capacity, because production hardware, traffic mix, and regional conditions can differ from the benchmark. They can postpone optimization when the test has at least 30% headroom, errors are near zero, soak results are stable, and the cost model remains acceptable. Capacity planning is an ongoing release activity, not a single prelaunch milestone.
A Defensible NGO Load-Test Recommendation for 2026
Start by writing one sentence that defines supported load: for example, “Release candidate 1.4 sustains 500 clients across 10 instances of 50 players, with p95 server tick time below 20 ms, p95 connection time below 5 seconds, unexpected disconnect rate below 0.1%, and no memory growth above 5% during a four-hour soak.” The exact values should reflect the game, but this form is useful because every term can be measured. Establish those thresholds before tuning hardware, then run smoke, ramp, expected-peak, stress, and soak profiles. Repeat important runs at least three times, retain raw data, and compare packaged clients with the release candidate server. A target of 500 clients should never be reported as achieved if only 100 clients were injected while 400 simply waited for slots.
For smaller indie teams, begin with Unity’s supported multiplayer-player testing features, a small number of dedicated Linux servers, and a scripted bot scenario rather than purchasing a large platform immediately. This approach can reveal replication mistakes and basic CPU limits with modest expense. Before a closed beta, public launch, seasonal event, or major content release, move the same scenario to a managed or self-hosted fleet that resembles production and reaches at least 110% of forecast peak. A 20% over-target stress test is a reasonable initial margin, though high-growth or event-driven products may need more. Hold the test at target load for several hours and repeat it after NGO, Unity, server, transport, or build changes; major networking changes warrant complete reruns.
The final decision should combine performance, cost, and operational confidence. If 1,000 clients cost nearly twice as much as 500 while meeting the same experience target, redesign or optimize; if the added capacity costs little and preserves headroom, scaling may be sensible. If soak testing reveals gradual degradation, do not launch regardless of the impressive peak result. Record the result as a tested configuration, including date, commit, package versions, scenario, region, and instance type. On September 29, 2026, the most authoritative statement is not a universal NGO player limit, but the highest workload your exact build and hosting topology have demonstrated under stated conditions. That evidence can guide a studio’s tools budget, multiplayer operations plan, and launch commitment without pretending that a benchmark removes uncertainty.