| Takeaway | Detail |
|---|---|
| Frame tails get a Boolean gate. | At 60 FPS, LaptopJudge gives an explicit 16.67 ms reference; at 144 FPS, the reference is about 6.94 ms. PASS requires steady frame times and trusted-logger inspection of 1% lows and spikes, not an average-FPS headline. |
| Interaction authority gets its own gate. | LaptopJudge targets idle jitter under 5 ms and 0% packet loss, with readings above 10 ms as a warning sign. A 40 ms local frame can feel like network lag even when ping is stable, so local and network failures must remain separate. |
| Load pressure cannot inherit a PASS. | Record a clean baseline before settings change, capturing ping, jitter, packet loss, frame stability, and temperature in the same session. LaptopJudge gives below about 85°C as a practical CPU temperature condition; missing memory-pressure evidence still fails closed. |
| The slowest signed artifact decides. | Use ping -n 100 8.8.8.8, record minimum, maximum, and average response, and probe the router and, where possible, game server. The final envelope comes from the slowest signed artifact; missing or contradictory evidence means HOLD. |
At 60 Hz, LaptopJudge's 2026 frame-time reference is 16.67 milliseconds. Three 30-minute runs contain 324,000 display deadlines that cannot be averaged away. The release rule is four Boolean checks: SHIP only when every gate passes, otherwise HOLD. Authority sits with the slowest signed artifact, not the faster editor profiler, because a quick capture cannot overrule the build players receive.
Frame stability is check one. LaptopJudge lists about 6.94 milliseconds at 144 FPS and recommends steady frame times rather than a headline average; CapFrameX or PresentMon can expose 1% lows and spikes. Check two treats interaction authority separately: idle jitter targets below 5 milliseconds, readings above 10 milliseconds are warning signs, and packet loss targets 0%. A stable ping cannot excuse a local 40-millisecond frame, nor network delay be mislabeled as frame time.
Check three requires a clean, comparable load baseline: record average ping, jitter, packet loss, frame stability, and temperature in the same session, with CPU temperature below about 85°C when practical. Check four demands signed-artifact attribution: send 100 test packets with ping -n 100 8.8.8.8, record minimum, maximum, and average response, then probe the router and, where possible, the game server. The envelope fails closed when evidence is missing or contradictory; frame tails, authority lag, and memory pressure remain independent blockers.

Present Deadline
Average FPS is the wrong control variable because it can hide recurring deadline loss and one catastrophic compositor stall. On the slowest certified 60 Hz target, calculate the deadline as 1,000 ms divided by Hz, record an OS present-to-present interval for every compositor-presented frame, and classify every interval against that deadline. At 60 Hz, apply the canonical 16.7 ms limit while retaining full measurement precision. Mean render duration is not a substitute: it omits time outside the renderer and between compositor presents.
Retain the complete interval stream and compute both p99 and the maximum present-to-present interval for every run. The p99 exposes recurring missed deadlines; the maximum prevents a catastrophic hitch from disappearing inside an aggregate. Keep p50 for diagnostics only—a healthy median cannot offset a failed tail. Do not average across runs or use a healthy run to dilute a failed one.
Measure response on two independent clocks. For local response, pair the input-event timestamp with the photon timestamp. For multiplayer state, measure the age of the newest authoritative snapshot at the client observation point. Ping, render duration, and server CPU time help diagnose neighboring parts of the pipeline, but none substitutes for either clock. Evaluate the two percentiles independently because acceptable local response does not prove current multiplayer state, and current state does not prove responsive input.
Sample process working set once per second, retain the engine allocation high-water mark, capture graphics-memory usage, and preserve trim, OOM, crash, ANR, and eviction evidence. On unified-memory devices, publish one combined physical footprint plus a CPU/GPU breakdown with non-overlapping allocation ranges; summing both views can double-count shared allocations. Every applicable CPU, GPU, or unified-memory peak must remain at or below 85% of its certified budget.
Use Unreal Insights in a separate diagnostic pass to attribute CPU, GPU, and allocation stalls. Run the ship gate itself with low-overhead release telemetry. A full profiler capture can change scheduling and alter the frame being diagnosed, making it useful for attribution but unsuitable as the authoritative qualification measurement.
Automate a fixed sequence: lock and verify refresh mode, warm the content, then execute three 30-minute post-warm-up release-build runs using the same scene, input script, and network fixture. Export one schema-validated, machine-readable record per run containing the device, OS, build hash, present percentiles and maximum, both lag percentiles, memory peaks, pressure events, and a trace ID. That ID should correlate the lightweight gate record with any later diagnostic trace; it does not imply that a heavyweight profiler was attached. An unlocked refresh mode, stale build identity, absent telemetry, or missing required field fails closed.
| Gate signal | Required observation | Pass condition for every run | Decision |
|---|---|---|---|
| Present cadence | OS interval for every compositor-presented frame | p99 at or below 16.7 ms; maximum at or below 33.3 ms | Any exceedance: HOLD |
| Local response | Input-event timestamp to photon | p95 at or below 50 ms | Failure or missing data: HOLD |
| Network state | Newest-authoritative-snapshot age under the recorded reference profile | p95 at or below 2.5 target server ticks | Failure or missing data: HOLD |
| Memory | One-second working-set samples, allocation high-water mark, graphics usage, and unified-memory reconciliation | Every applicable peak at or below 85% of its certified budget | Any breach: HOLD |
| Stability | Crash, ANR, OOM, trim, and eviction event log | Every count equals zero | Any nonzero or unavailable count: HOLD |
| Run set and record integrity | Three complete 30-minute records with every required field and trace ID | All prior conditions are fully green in all three runs | SHIP only when complete; otherwise HOLD |

Published Evidence
For a 2026 release decision, the published record makes a SHIP call harder to fake: each source exposes a different blind spot that an aggregate “looks smooth” check can miss. The evidence standard should therefore be a warmed release build, instrumented throughout the prescribed gate on the slowest certified target—not a favorable editor snapshot.
| Evidence area | Published figure | Decision consequence |
|---|---|---|
| Baseline profiles | According to Google’s Android Developers Baseline profiles documentation, improvements reach up to 40% for cold-start execution, 20% for warm-start execution, and 35% for janky-frame reduction. | Treat these as reported upper bounds, not guaranteed gains. Warm the build before the steady-state gate; startup improvement is not evidence of sustained release behavior. |
| App-not-responding detection | Google’s Android Developers guidance identifies 5 seconds of unresponsiveness as user-visible failure. At 60Hz, that span is approximately 300 frame intervals. | An ANR-only check can miss hundreds of frame deadlines before failing, so direct frame and latency evidence must accompany the failure ledger. |
| Low-memory Android tier | Google’s Android 14 memory documentation caps the foreground app heap at 256MB on devices with 512MB of RAM or less. | Use that low-end tier to require per-device budgets. The heap cap is neither a total process-memory ceiling nor a graphics-memory ceiling. |
| Unity frame capture | The Unity Scripting API documentation for FrameTimingManager.GetLatestTimings exposes no more than four recent frame timings per capture. |
A short editor snapshot cannot support a long-run percentile. Use sustained release-build capture rather than treating a small latest-frame sample as representative. |
| Refresh-dependent budgets | According to Meta’s Quest performance guidance, frame budgets are 13.9ms at 72Hz, 11.1ms at 90Hz, and 8.3ms at 120Hz. | Those figures establish that 16.7ms at 60Hz is a target-specific policy, not a refresh-independent quality fact. |
These mechanisms do not compete. Baseline profiles describe potential execution benefits; Android’s ANR threshold defines a coarse watchdog boundary; its heap ceiling has a deliberately limited scope; Unity’s capture API has finite sample depth; and Quest budgets change with display cadence. Combining the sources does not create a substitute for measurement—it identifies what an apparently healthy summary may fail to cover.
A clean aggregate frame trace therefore cannot authenticate uncaptured samples or replace the other evidence streams. It may coexist with poor input-to-photon lag, stale multiplayer state, or a memory-budget breach. The practical skill is evidence triangulation: retain per-run traces, maxima, percentiles, memory peaks, and event logs under the recorded reference profile rather than accepting a summary screenshot as proof.
Apply the decision rule without exception: SHIP only when every prescribed post-warm-up release-build run is complete and fully green; a failed or missing criterion yields HOLD. Published upper bounds and platform limits inform test design, but they never waive a release-build observation.

Gate Comparison
The release decision should be a four-boolean record, not a score. Frame tail, interaction plus authority lag, memory, and stability form independent gates. The lag gate requires both perceived-interaction and authoritative-snapshot criteria; the memory gate requires every applicable budget check; the stability gate rejects crashes, application-not-responding events, OOMs, trims, or evictions. Overall SHIP eligibility is the logical AND across every required complete run. Any false—or missing—observation yields HOLD. An average frame rate, a clean sub-deadline graph, or an excellent input result cannot cancel a missed frame, stale authority, memory breach, or stability event.
| Method | Frame truth | Lag truth | Memory truth | Verdict |
|---|---|---|---|---|
| Editor profiler | Detailed CPU/GPU, but wrong runtime | Profiler changes timing | Allocation capture can perturb measurement | DIAGNOSTIC ONLY |
| Generic cloud device | Repeatable only when hardware and display image match | Route and display vary | Memory model may differ | TRIAGE ONLY |
| Hybrid on-device hardware-in-the-loop release trace plus sampled camera | Exact signed build and OS deadlines | Instrumented input plus photon calibration | OS footprint and graphics high-water in the same run | WINNER |
| Live cohort telemetry | Real-world distribution | Field input and network | Field crashes and pressure | POST-SHIP CANARY |
The hybrid on-device hardware-in-the-loop release trace is the explicit winner. Among these candidates, it alone can bind frame-deadline behavior, perceived lag, and memory pressure to the same signed artifact, slowest certified target class, and run. That co-location makes a breach diagnosable rather than an artifact of measurements stitched across different sessions. On consoles and handhelds, use the platform-native equivalent: signed-build tracing, vendor deadline telemetry, calibrated input-to-photon capture, and OS graphics-memory instrumentation.
A deterministic PR smoke scene remains useful for catching catastrophic regressions before a release candidate advances. Its deterministic workload and repeatable input make failures easier to localize, but a smoke pass cannot certify the candidate’s full runtime behavior. Reserve the complete hybrid gate for the release candidate, and use live-service telemetry afterward as a canary. Neither phase may substitute for the release gate or convert missing evidence into a pass.
Finally, calibrate tracing overhead on the same build, device, workload, and reference profile used for the gate. Compare traced p99 present time with a low-overhead run; calculate the relative shift as the traced-minus-baseline difference divided by the baseline. If tracing shifts p99 by more than the stipulated 5%, reject that measurement as gate-authoritative. The attempted run then cannot support SHIP: use a lighter capture or rerun the complete gate, and return HOLD if authoritative evidence remains unavailable.

What the Data Doesn't Tell You
A clean dashboard can still be inconclusive. The complete release-build gate on the slowest certified 60Hz target is a pass/fail artifact, not a general warranty of smoothness, visual quality, thermal stability, memory safety, or network behavior. Its limits matter: under the canonical rule, uncertainty does not create a softer status. A failed—or unrecorded—criterion remains HOLD.
The following are stipulated counterexamples, not claimed field measurements:
| Evidence artifact | What it appears to show | What it does not establish | Required release treatment |
|---|---|---|---|
| Average-FPS trace | 98 present intervals of 16.0ms plus two 40ms intervals average 16.48ms, or about 60.7 FPS. | Release-safe timing; p99 lands in the 40ms stall tail. | Calculate p99 within each run; never substitute average FPS. |
| Resolution change | Switching from 1080p to 720p removes 55.6% of rendered pixels while leaving the timing trace green. | Equivalent visual quality or universal genre suitability. | Log resolution scale and treat declared thresholds as release policy, not visual-quality guarantees. |
| Thermal soak | A 30-minute run never enters a vendor-defined severe thermal state. | Behavior beyond that window or on another device. | Record thermal status throughout; extend testing to 60 minutes for heat-sensitive hardware. |
| Memory trace | A 220MiB steady CPU footprint coexists with a 200MiB graphics burst immediately before a 32ms render-thread stall. | That sparse averages represent transient allocation pressure. | Capture allocation high-water marks and per-frame maximums. |
| Network fixture | A fixed route passes while a congested mobile client encounters 80ms RTT and 1% packet loss. | Success across the live player population. | Record the reference profile and keep the network gate explicitly tied to that profile. |
| Pooled frame samples | Combining frames from three runs produces one tidy percentile. | That each run met its tail limit; correlation can conceal a bad run. | Calculate p99 per run and report the worst-run margin, not a pooled percentile. |
The average is arithmetic, not temporal evidence. It compresses the two longest present intervals into the same headline that describes the other 98, so a roughly 60 FPS summary can coexist with a visible stall tail. The correct unit of judgment is the individual run: a pooled distribution can improve merely because healthier samples outweigh a failing run.
Timing and thermal results also depend on declared conditions. Reducing rendered pixels may satisfy the timing policy while changing the presentation workload, which is why resolution scale belongs beside the trace. Likewise, a vendor-defined thermal state is meaningful only for the recorded device, status history, and test duration; absence of a severe state during the required soak cannot be projected indefinitely.
Memory and network summaries fail in opposite directions. Sparse memory sampling can hide a short coupled allocation-and-render event, while an over-clean network fixture can hide adverse client conditions. High-water marks, per-frame maxima, and an explicit reference profile preserve those distinctions. These limitations do not soften the decision: only three fully green runs support SHIP; any failed or missing criterion leaves the release on HOLD.

Worked Gate
HOLD is the only defensible release label when the three measured trace packages are absent. The supplied material contains no release-build hashes, raw compositor intervals, calibrated photon events, authoritative-snapshot timestamps, memory series, or trace locations. Those omissions are not neutral: under the gate, missing evidence is failure; inventing plausible rows would turn an audit into marketing.
I would commit-pin a fork of godotengine/godot-demo-projects’ 3D Platformer, export Godot 4.5 for iOS 26 with ENetMultiplayerPeer, and treat the iPhone 13 as this case’s slowest certified target. The tested release artifact must remain locked and verified at 60Hz, with editor and debug overlays removed. Each run uses one physical client, seven headless clients, and a dedicated server on the same LAN. After a five-minute warm-up, record 30 measured minutes while a repeatable script exercises movement, combat, asset streaming, joins, and reconnects. Linux tc netem then imposes the recorded reference profile: 40ms RTT, 5ms jitter, and 0% loss. Archive the qdisc configuration and node roles with the traces.
The case’s certified unified-memory budget is 640MiB, producing a 544MiB gate at 85%. A 30Hz server tick makes 2.5 ticks an 83.3ms snapshot-age ceiling. Sample phys_footprint once per second. Keep the graphics-allocation high-water mark as a separate diagnostic; adding overlapping allocation counters would double-count memory and invalidate the result.
Capture every compositor present-to-present interval. For each run, timestamp at least 1,000 input flashes with a 240fps camera having 4.17ms temporal resolution. Publish the calibration error and derive input-to-photon p95 from captured photons, not render completion. This separation is what keeps a clean frame graph from concealing interaction lag.
No actual trace data was supplied with this section, so the evidence-admission result must remain explicit rather than being populated with illustrative numbers:
| Run | Build hash | Frame p99 / maximum | Input-to-photon p95 | Newest-authoritative-snapshot p95 | Peak unified memory | Crash / ANR / OOM / trim / eviction | Trace ID or URL | Decision / exact worst margin |
|---|---|---|---|---|---|---|---|---|
| 1 | Not supplied | Missing / missing | Missing | Missing | Missing | Not reported | Not supplied | HOLD / undefined: raw operands missing |
| 2 | Not supplied | Missing / missing | Missing | Missing | Missing | Not reported | Not supplied | HOLD / undefined: raw operands missing |
| 3 | Not supplied | Missing / missing | Missing | Missing | Missing | Not reported | Not supplied | HOLD / undefined: raw operands missing |
A row earns SHIP only when frame p99 is at or below 16.7ms, maximum is at or below 33.3ms, input-to-photon p95 is at or below 50ms, snapshot-age p95 is at or below 83.3ms, peak unified memory is at or below 544MiB, and every failure count equals zero. Calculate signed slack from unrounded values as limit minus observation, retaining milliseconds or MiB; round only displayed values. Because no raw operands exist here, the exact worst margin is undefined—not zero. Overall result: HOLD, with zero complete green rows out of the required three. No 60 FPS average or sub-16.7ms frame graph can promote missing or failing input, authority, memory, or stability evidence.

How to Choose Well
Choose by conjunction, not by a composite score. A healthy average frame rate—or a graph that appears under deadline—does not prove ship readiness: tail misses, input-to-photon lag, stale multiplayer state, and memory pressure are independent vetoes. Under the specified release rule, publish one binary label with its trace package, not a grade that lets strength in one area conceal failure in another.
Freeze hardware and mode identity before testing. Lock the slowest certified target to the prescribed base-mode gate and preserve the budget definition used for that target. If the release advertises another refresh mode, require a separate gate at that mode’s own deadline; never inherit a pass from the base mode. Any supported but unproven mode leaves the release HOLD.
Read frame evidence per run. The tail percentile exposes recurring budget loss; the longest present interval exposes a one-off compositor, allocation, or recovery stall. A green frame label requires both checks, but it authorizes only that gate—it cannot compensate for lag, memory, crash, ANR, or telemetry failure. Smoothing or clipping the trace destroys the evidence needed to distinguish those cases.
Keep perceived response separate from network authority. Match input and photon timestamps for the input-to-photon distribution rather than substituting frame cadence. For networking, independently timestamp the newest authoritative snapshot and compare its age with the declared tick contract while the recorded reference-network profile is active. A breach in either distribution is HOLD: responsive local rendering can still conceal delayed state.
Treat memory as a veto, not a score. Compare every applicable CPU and graphics peak with the certified budget for that pool, and treat OOM, trim, and eviction log entries as failures rather than recoverable telemetry. On unified memory, test the combined certified footprint once; adding CPU and graphics allocations again invents pressure the platform does not carry and can falsely block an otherwise valid build.
Join the results with logical AND, never an average. Every required release-build run must be complete, post-warm-up, fully instrumented, and clear of the frame, lag, memory, crash, and ANR gates. A missing trace package is a missing result, not permission to infer a pass. One failed or absent run therefore yields HOLD; SHIP exists only when every required run is green.
| Decision rule | Required condition | Decision |
|---|---|---|
| 1. Mode | Is the slowest certified target locked to 60Hz, and is every other advertised mode separately proven at its own deadline? | Any unproven supported mode: HOLD. Otherwise, continue. |
| 2. Frame | For each of the 3 runs, is p99 present-to-present time at or below 16.7ms, with no present interval above 33.3ms? | Any breach: HOLD. All pass: mark only the frame gate green. |
| 3. Lag | Is p95 input-to-photon at or below 50ms and, if networked, p95 newest-authoritative-snapshot age at or below 2.5 target server ticks under the recorded reference-network profile? | Either breach: HOLD. Both applicable checks pass: continue. |
| 4. Memory | Is every applicable CPU or GPU peak at or below 85% of its certified budget, with OOM, trim, and eviction counts at zero; is unified memory tested once at 85% of the combined footprint? | Any breach or event: HOLD. Otherwise, continue. |
| 5. Release | After the mode check, do all 3 complete, post-warm-up 30-minute release-build runs clear every applicable gate with complete telemetry and crash and ANR counts at zero? | SHIP only when complete; otherwise HOLD |
Frequently Asked Questions
What present-to-present limits does the 60 Hz ship gate enforce, and what reference is given for 144 FPS?
The 60 Hz gate requires p99 at or below 16.7 ms and a maximum at or below 33.3 ms, while the 144 FPS reference is about 6.94 ms.
Can a healthy median or one successful run hide missed deadlines?
No: p50 is diagnostic only, and all three complete 30-minute post-warm-up release-build runs—324,000 display deadlines total—must keep every gate green without averaging or dilution.
Can acceptable local input response certify that multiplayer state is current?
No: local input-event-to-photon p95 must be at or below 50 ms, while newest-authoritative-snapshot age p95 must be at or below 2.5 target server ticks under the recorded reference profile, measured on two independent clocks.
Does stable ping excuse a 40 ms local frame or nonzero packet loss?
No: local and network failures remain separate, idle jitter targets below 5 ms, readings above 10 ms are warning signs, and packet loss targets 0%.
What memory and failure-ledger values are required for a SHIP decision?
Every applicable CPU, GPU, or unified-memory peak must be at or below 85% of its certified budget, and every crash, ANR, OOM, trim, and eviction count must equal zero.
What network evidence is required for the slowest signed artifact's final envelope?
Run ping -n 100 8.8.8.8, record minimum, maximum, and average response, probe the router and, where possible, the game server, and treat missing or contradictory evidence as HOLD.
Quick answers
| What is the release rule regarding the four Boolean checks? | SHIP only when every gate passes, otherwise HOLD. |
| What does Check one (frame stability) require according to LaptopJudge? | Steady frame times and trusted-logger inspection of 1% lows and spikes, not an average-FPS headline. |
| What are the targets for Check two (interaction authority) regarding idle jitter and packet loss? | Idle jitter under 5 ms and 0% packet loss, with readings above 10 ms as a warning sign. |
| What must be recorded in the same session for Check three (load pressure baseline)? | Ping, jitter, packet loss, frame stability, and temperature. |
| What action is required for Check four (signed-artifact attribution) involving network testing? | Use ping -n 100 8.8.8.8, record minimum, maximum, and average response, and probe the router and, where possible, game server. |
Also worth reading: Rollback Netcode: Free GGPO, Photon Fusion, and One Winner: Rollback Netcode: Free GGPO, Photon · Why Rainbow Six Siege Has No Pacifist Option: Code and Data: Why Rainbow Six Siege Has · iPhone Controller Compatibility: Test Xbox and DualSense on 14.5 vs Current: iPhone Controller Compatibility: Test Xbox