The action is the binding.
Plans, model/profile identity and current context travel together. A modified suffix, changed scene or stale observation cannot inherit an unrelated verdict.
ROBOT POLICY VERIFICATION · OPEN RESEARCH
An execution gate for transformed robot plans, with single-use permits and independently verifiable evidence.
Sentinel EVC Lab · LancerLSY
01 / DEMONSTRATION
02 / OVERVIEW
Robot policies produce action blocks that may be retimed, repaired, transformed between frames or spliced before execution. A verdict on the original plan does not automatically cover the resulting motion. Sentinel EVC binds the verification decision to the final action and current execution context.
The local research workbench combines final-plan checks, bounded certificate reuse with full fallback, a single writer to the controller and signed evidence. Separate research profiles evaluate UR5e motion, WorldGuard consequence prediction and SmolVLA action reconstruction on recorded SO100 data.
03 / METHOD
Plans, model/profile identity and current context travel together. A modified suffix, changed scene or stale observation cannot inherit an unrelated verdict.
The UR5e study reuses only a matching static prefix. Dynamic rollout always starts from frame zero; an invalid parent record triggers full validation.
Submitted, accepted and observed commands are distinct. Revocation blocks new old-generation submissions; recorded events and digests support independent verification.
The product core supports bounded numeric and fixed contact profiles. Trained neural models and the UR5e study are separate research profiles; a live robot motion adapter is not enabled.
Read the mechanism and design trace04 / RECORDED 3D REPLAY
Follow the proposed final motion, then compare the two execution decisions.
Original UR5e visual meshes, with forward-kinematic poses from 41 saved joint states per case. No new dynamics or interpolated poses. Rejected motion is shown as a counterfactual, not dispatched motion. The examples were selected after review; collision labels describe the entire trajectory, not individual frames. The blue path follows Wrist 3, not a mounted tool. UR5e model: MuJoCo Menagerie, BSD-3-Clause.
05 / EXPERIMENTS
Select a research figure to open its full-size version. Counts and evidence scope are explained below each figure.
The protocol and source hashes were committed before formal runs. Under a declared friction floor, the long action completes 100/100 low-friction tasks; camera mismatch routes to the camera-independent state model. Unsupported mass and goal profiles still reject 300/600 roots.
The router completes 300/600; fixed 4.8 s completes 600/600. Rejection is incomplete. Zero unsafe observations per 100-root scene has a 3.70% Wilson upper bound. Trusted profile declarations do not measure friction.
A separately frozen 27-cell stress study evaluates 324 new roots. Fixed 4.8 s completes 324/324 with no observed unsafe/drop events; 1.6 s completes 184/324 and drops 78. The floor rule equals fixed 4.8 s on every root, at 3× duration. Descriptive 95% Wilson zero-event upper bounds: aggregate 1.17%; each cell n=12: 24.25%. This does not qualify the continuous parameter domain.
On 60 predefined UR5e roots, with obstacles independently sampled before plans, full and incremental decisions agree. Mean marginal saving is 0.996 ms; incremental P95 is worse. No tail-latency gain is claimed.
Frozen design, failures, paired results and intervalsSix scenario classes compare four policies on the same roots. A separate reviewer uses denser static sampling and 1 ms MuJoCo replay without importing the gate runner.
106 static-prefix reuses and 74 full fallbacks. Incremental validation costs 147.65 ms including the parent, versus 48.74 ms for full validation: no speedup claim. These are fixed-order observations on constructed cases, not a natural-distribution holdout or continuous collision proof.
Methods, denominators and cost accountingA 100-rollout confirmation, with source and protocol frozen before execution, succeeds in 93/100 (descriptive 95% Wilson interval 86.25–96.57%), with zero crashes. Seven 280-step failures remain: four in the top-drawer task, one each in tasks 3, 5 and 8. All 11,592 official step outcomes and unchanged environment actions were independently audited. This evaluates the official native Panda checkpoint, separately from the SO100 overlay and Sentinel intervention benefit.
Under MuJoCo 3.8.1 with a 50-action execution horizon, the official SmolVLA LIBERO checkpoint succeeds in 58/100 fixed rollouts (descriptive 95% Wilson interval 48.21–67.20%). The bowl-on-ramekin task fails 10/10; all 42 failures reach the 280-step cap. There are 17,868 unchanged postprocessor-to-env.step actions. This measures the native official checkpoint and logging path, separately from the SO100 overlay and intervention benefit.
40 no-policy resets isolate MuJoCo 3.8.1 versus 3.3.7 with unchanged task files, assets and paired seeds. Task 5 bowl position differs by 35.327 mm in all 10 pairs; control mean shift is 0.688 mm. This confirms version-sensitive initial conditions, without establishing why the policy failed. The older engine is a benchmark-compatibility comparison, not more accurate physics.
A frozen vision-language backbone with a fine-tuned action expert. The pinned base and development-selected overlay use identical sampling noise in one paired held-out evaluation.
Action reconstruction, not physical robot task success. Windows overlap; the independent unit is the episode. Dataset-native physical units are undeclared.
Action MAE / training action standard deviation
63.46% lower normalized MAE
600 fresh roots, six scenarios, four sibling plans per root. The trained state and visual ensembles keep their original calibration; difficult cases remain in the comparison.
Unsafe selections / 100 roots · lower is better
Full scenario results and uncertaintyMUJOCO PHYSICS-PROFILE FALLBACK
A separate 5 s MuJoCo profile evaluates 100 new low-friction roots. Its 4.8 s candidate has 0/100 unsafe selections and 100/100 tray endpoint tasks, at three times the 1.6 s duration.
The rule assumes μ_min = 0.015; friction is not sensed. The 3.2 s candidate was also safe in this sample but rejected by the conservative rule. Zero observed failures has a 3.70% Wilson upper bound. This profile is not deployed and does not repair the original neural gate.
Fallback rule and retained resultsGPU measurements: RTX 4090 D (24 GB). Every chart uses retained result files, available with source hashes in the evidence download.
Download chart data + source identities06 / RESEARCH WORKBENCH

The installable CLI and macOS App share a local engine. Run numeric and fixed MuJoCo contact profiles, replay recorded trajectories and export signed evidence bundles.
Research prototype. The macOS App is locally built without distribution signing or notarization; imported models do not acquire motion permission. The physical robot port remains read-only.
Installation and supported capabilities07 / REPRODUCE
git clone https://github.com/LancerLSY/sentinel-evc-lab.git
cd sentinel-evc-lab
python3 -m venv .venv
source .venv/bin/activate
python -m pip install -e ".[physics]"
python -m sentinel_evc serveStart from source with Python 3.10+, or use the interactive installer to choose a core or physics profile and build the macOS App. The CLI prints the local workbench address; simulations write recorded outcomes and evidence.
GPU training uses a separate pinned environment. Model overlays require the original frozen SmolVLA base/backbone; they are not standalone policy checkpoints.
Training, calibration and evaluation commandsThe historical simulator directory checksum includes download cache. Diagnostic replay source binds official model content per file. Replay reproducibility note
The release index includes sizes and SHA-256 digests; archives contain per-file identities and licenses. Sentinel code is MIT; MIT does not grant patent rights. See the repository for upstream attribution and the core mechanism's patent notice.
Asset index and checksums08 / CITATION
Use the repository citation for the current software and simulation artifacts. A paper citation will be added alongside its arXiv record.
@misc{sentinel_evc_lab_2026,
title = {Sentinel EVC Lab},
year = {2026},
url = {https://github.com/LancerLSY/sentinel-evc-lab},
note = {Research software and simulation evidence}
}