Follow the result beyond the ACK.
Isolated physical executions · same action, controlled decision delay
p50 and p95 include the decision and downstream work. Decision shares are the reported mean shares across scenes.
A decision has to be correct, arrive in time, and produce the intended result. Explore where faster interpretation helps wireless control and edge services—and what it takes to verify the loop.
Explore the evidenceJev calls fit a 1 s timing reference.
Interpretation + fixed 5 ms E2 · correctness and radio outcomes explored below02 · Edge orchestration90.7–95.3%Correct, on-time service across 1–16 req/s.
Jev · live admission + modeled execution · four interpretation slots03 · Closed-loop verification100 → 0%The same decision can pass alone and fail in a queue.
Transport · 0.5 s decision · 10 s budget · measured execution replayWhat reaches the radio, and what changes for users? Read the service outcome alongside policy quality, control deadlines and queues.
Affected-class SLA headroom relative to no update.
95% interval 2.17–7.02 pp · base pointHigher violation with 5 s instead of 0.1 s enforcement.
95% interval 1.22–6.37 pp · true policy held fixedFaster interpretation reduces the time spent waiting for a policy. The base-point affected-class SLA comparisons do not establish a hosted-model ranking under both block lengths.
Latency-only ns-3 closed loop · event-driven mode · 2 s SLA window
All models are shown together. Select a measured rate and UE speed.
Select a cell to inspect all four KPIs. Arrow keys move between cells; on small screens, swipe the map horizontally.
Link failures are counts per run; run durations differ across arrival rates. These radio KPI point estimates are descriptive.
57-cell telemetry · matched quality and timing conditions
Full-policy correctness and timing fit are separate marginal rates from the same condition. Their product is not a measured joint success rate. Timing uses interpretation + fixed 5 ms E2 within 1 s.
For fresh 57-cell telemetry, Jev's API fee is $0.207 per 1,000 correct policies, the lowest among the evaluated hosted interpreters.
4,800 calls per model · pooled over 16 telemetry conditions
Drag the slider or click the plot to select a budget. Hover previews values; moving away restores your selection.
Set a target, then select a model's first qualifying sampled budget.
This lookup uses the 61 sampled budgets from 10 ms to 10 s, with no interpolation. Passing means timely, error-free interpretation plus the fixed E2 term; policy correctness and full-loop success are separate.
A miss is an errored call or interpretation latency + 5 ms above τ. Architectural near-RT/non-RT bands describe control-loop timescales; the whole budget is not reserved for the interpreter.
Four interpretation slots · 300 intents per model and rate
η is utilization across all four slots. Values above 1 describe a non-stationary measurement window. Each rate uses a separate live trace: hosted response times also vary between traces, so these points are not a controlled latency-only capacity curve.
srsRAN gNB · O-RAN SC near-RT RIC · A1 simulator · 30 arrivals/model
Stacks sum component medians. KPM is a separate observation upper bound and includes its reporting window, convergence and traffic ramp-up. It is not an additional segment to add to the ACK stack or a verified recovery time.
Radio outcomes: RQ1–2, latency-only ns-3 runs. Base-point SLA intervals use 48.8 s blocks, with 10 s sensitivity checks; grid points are descriptive. Policy quality: RQ4, 57-cell telemetry. Deadline curve: supplied empirical aggregates at 61 budgets; all seven 1 s values match the manuscript, without a fresh raw-call reanalysis. Load: RQ3, live arrivals and modeled enforcement. Control path: RQ6, real-stack component medians and separately observed KPM bounds.
How many requests finish correctly before their deadline? Keep completion and latency together, then inspect what limits a real service.
A stays pinned while the controls above update B.
○ A · ● B · horizontal position is correct, on-time completion (0–100%). Differences are descriptive.
Request p95 belongs to each condition's own completions. A lower p95 with fewer completions does not establish a latency advantage.
Read the two outcomes together. Request p95 is conditional on each model's own completions. A low p95 with few completions does not establish a latency advantage. The count column reports the selected correct-completion criterion.
Part A uses live admission and modeled execution. Completion requires the full intent to be exactly correct and the supported service to finish on time (300 supported requests per cell). Part B uses real OCR: correct completion additionally requires the recognized text to match (180 supported OCR requests per cell). Its 60 unsupported requests and correct rejections are counted separately. Counts are recovered from manuscript-rounded rates using these fixed denominators and checked by rounding back. Request p95 is computed on each model's own completions; “—” means no completed requests. Supporting study: RQ5.
Where does the clock stop? What happens when work queues? Who checks an infeasible action? Measured cases show why these choices change the verdict.
The guide reads reported aggregates; it does not animate an individual execution trace. Checkpoints 1–3 describe the isolated arm. Checkpoint 4 changes the test context to the selected queue and replay arm.
Isolated physical executions · same action, controlled decision delay
p50 and p95 include the decision and downstream work. Decision shares are the reported mean shares across scenes.
Real single-slot FIFO queue + measured execution-time replay. Rates stay fixed across delay arms. Intervals describe the finite measurement window, including overloaded queues.
Jev-1.13 · constructed application-control tasks · different checks retain their own denominators
Requests a new candidate when the catalogue has no valid option.
Escalates when the forwarding instance is infeasible.
The workflow intervention changes instructions, refresh actions and check ownership together. Full rule: 120/120 in both arms for these quota conditions.
Named endpoints · correctness and failure counts · deadline attainment · arrivals and concurrency · queueing · check ownership · verified outcome
Of 50 families claiming a control-loop or timing fit, four provide matched measurement. Across all 139 families, nine report p95 or higher and four report deadline attainment.
Read the coding uncertainty. Load, queueing, stability, network round-trip inclusion and final-state checks have lower pre-adjudication coding agreement. The family bootstrap intervals below do not capture that coding uncertainty. Categories overlap and all 139 families remain in the denominator.
S = selection · G = generation · C = deterministic computation · U = unspecified. “Unresolved” means the map does not establish check ownership. Check descriptions support inspection, not a reliable prevalence estimate or a deployment ranking.
Fixed-action endpoint measurements use 30 scenes × 10 repeats per delay and domain. Load tests use the same decision queue with execution residuals sampled from measured physical runs; the post-queue physical actions do not run concurrently. Transport and edge retain their own exploratory 10 s and 2 s budgets. Correctness checks are separate constructed tasks. Reporting coverage and the 139-family map come from the current SoK tables and supplement. Supporting study: diagnostic lenses L1–L3 and reporting coverage.
Each direction links to the study that supplies its measurements, methods and evidence.
Delong Li, Xu Wang, Haochen Gong, Rui Lang, Guangsheng Yu
Delong Li, Xu Wang, Haochen Gong, Rui Lang, Guangsheng Yu
Delong Li, Chen Li, Xu Wang, Haochen Gong, Rui Lang, Guangsheng Yu
Explore recorded experimental conditions. API fees refer to those measurements. Comparable energy per decision is unavailable across hosted and self-hosted deployments.
Download chart data ↓This is the same aggregate data embedded in the page. You can select and copy it if your browser does not support the download.