Control
0.2466
Single-turn baseline
CONEXUS reports controlled experiments and internal computational benchmarks with their conditions, baselines, methods, and source paths. Each finding is presented in its tested context.
Across 200 independent runs, the full CONEXUS architecture pushed the model farther from its ordinary response pattern than any of the three control conditions. More tokens alone did not reproduce the effect.
Four controlled conditions separated the complete architecture from an ordinary baseline, a longer neutral prompt, and token-only exposure. That comparison shows whether the measured change came from the CONEXUS sequence rather than prompt length or symbols alone.
What semantic distance means: Think of the model's answers as points on a map. Semantic distance measures how far those answers move away from the usual neighborhood. A higher score means the search reached farther into the measured idea space.
The study tested Gemini 3.1 Pro Preview on an Alternative Uses Task, with 50 independent runs per condition, temperature 0.7, 16,000 maximum output tokens, and local BGE embeddings for the semantic-distance measurement.
Control
0.2466
Single-turn baseline
Neutral
0.2219
Analytical multi-turn prompt
Token-only
0.2258
Emoji exposure without the architecture
CONEXUS
0.2929
Complete contradiction-holding sequence
d = 3.7824
This standardized effect-size measure describes how far apart the two conditions were.
Welch p-value: 2.97e-32
The bootstrap interval for the mean difference was [+0.063467, +0.078094].
p = 0.3612
Token-only and Neutral remained closely aligned in this comparison.
Remember the central result: the complete architecture changed the measured search pattern, while extra tokens alone did not.
The full CONEXUS architecture moved responses farthest from the model's ordinary response pattern.
Neutral and CONEXUS showed a very large separation in the tested configuration: d = 3.7824.
Token-only prompting did not reproduce the CONEXUS effect: p = 0.3612.
The longer neutral prompt did not expand the measured search behavior, so prompt length alone does not explain the result.
The Forgetting Engine tests whether better search can come from strategically removing low-value paths while preserving promising ones.
The idea was tested across different optimization problems. The locked optimization sweep contains 30,800 controlled trials, with additional domain studies in different search spaces. Each benchmark keeps its own objective, baseline, configuration, and measurement.
Benchmark context: protein folding reports a 561% relative success-rate difference, routing reports an 89.3% improvement, and quantum compilation reports a 27.8% gate reduction within their respective experiments.
2,000 trials
Approximately 80% relative improvement in the stated comparison
Internal benchmark against the documented Monte Carlo baseline
4,000 trials
25.8% success versus 3.9%, approximately 561% relative improvement
Largest reported relative gap in this research portfolio
Scale series trials
Larger relative gaps were reported at larger tested instances
Benchmark-specific trend across the tested scale series
250 trials
Up to 89.3% improvement at the largest tested scale
Compared with the stated routing baseline and configuration
300 trials
Reported accuracy gains ranged from 3.8% to 8.4%
Internal search benchmark under the documented setup
5,000 trials
27.8% gate reduction and 3.7% fidelity gain were reported
Simulator-based comparison under the documented compilation setup
In several CONEXUS benchmark series, the relative advantage over the chosen baseline increased at larger tested scales. CONEXUS calls that observed pattern complexity inversion.
The research record tracks that pattern across the documented benchmark families, baselines, objectives, trial counts, and tested scales.
Observed
Larger relative gaps in selected benchmark series as tested scale increased.
Research scope
Current evidence covers the documented benchmark families, baselines, objectives, trial counts, and tested scales.
This exploratory case study shows how strategic retention surfaced three anomalous signals in public catalog data and preserved them for follow-up astronomical review instead of eliminating them early.
Retained by the exploratory anomaly-ranking process for follow-up astronomical analysis.
Retained by the exploratory anomaly-ranking process for follow-up astronomical analysis.
Retained by the exploratory anomaly-ranking process for follow-up astronomical analysis.
The strategic-retention approach can surface and preserve anomalous candidates that might otherwise be eliminated early in a ranking pipeline.
Inspect the Four-Arm validation materials, the Forgetting Engine executive and full audits, and the research-validation repository directly.
CONEXUS will continue separating demonstrated results from research hypotheses, product concepts, and future applications.
Request Technical Materials