The problem
A camera can produce a consistent position estimate and still be wrong. I wanted to test whether a second, imperfect source of evidence could expose that error, and when the system should decline to continue.
The project started with synthetic landing experiments. Later evaluations introduced genuine Gazebo camera frames, partial views, appearance changes, and new geometry.
What I built
I built the simulation experiments, supervisory variants, evaluation protocols, and result archive. I compared image-only estimation with temporal checks, an independent estimate, uncertainty calibration, and abstention.
I froze each candidate before evaluating its holdout. That separation mattered: changing the model after seeing the test would make the same examples less useful as evidence.
The early result
On the Phase 6B synthetic benchmark, selective intervention reduced unsafe touchdowns from 43% to 1%, with a reported 3% timeout cost. That was evidence within a simplified experiment.
The next question was whether the approach would survive more difficult camera evidence. The answer was mixed.
The harder test
Phase 10R evaluated a frozen candidate on 36 sequences, covering 12 new geometry trajectories across three appearance conditions: 1,440 truth-visible frames.
Average error improved on ambiguous views. The candidate still failed the preregistered all-gates rule: tail error, missed observations, and uncertainty coverage did not meet their targets.
| Measure | Result | Gate |
|---|---|---|
| Ambiguous lateral mean-error improvement | 79.2% | Passed |
| Ambiguous altitude mean-error improvement | 73.7% | Passed |
| Truth-visible miss rate | 20.0% | Failed |
| Lateral 95% uncertainty coverage | 84.3% | Failed |
| Altitude 95% uncertainty coverage | 79.7% | Failed |
Both p95 improvement gates also failed. The complete protocol and result record are linked below.
What I learned
Improving the average did not solve the difficult failures. Calibration that looked good during development became overconfident after the appearance and geometry changed.
I kept the failed holdout result in the public record. My next question is whether an estimator can recognize when its uncertainty calibration has stopped transferring.
Scope
This is simulation research. I have not validated it on a physical aircraft or hardware camera. The Phase 10R holdout has been seen and cannot be reused as an unseen test.
See the ambiguity
This small educational model shows why a landing-marker estimate can become difficult when viewpoint and visibility change. Adjust the camera angle or cover part of the marker. The scene explains the geometry; it is separate from the measured Phase 10R results above.
Mostly visible: the marker outline is easy to interpret.
Teaching model only. It does not replay AegisLand measurements, replace a camera model, or change the published holdout result.
