CONEXUS homeView Company Site
Evidence and Methods

Measured results.Visible methods.

CONEXUS reports controlled experiments and internal computational benchmarks with their conditions, baselines, methods, and source paths. Each finding is presented in its tested context.

Primary Causal Study

We tested whether CONEXUS changes the way an AI searches for ideas. It did.

Across 200 independent runs, the full CONEXUS architecture pushed the model farther from its ordinary response pattern than any of the three control conditions. More tokens alone did not reproduce the effect.

Four controlled conditions separated the complete architecture from an ordinary baseline, a longer neutral prompt, and token-only exposure. That comparison shows whether the measured change came from the CONEXUS sequence rather than prompt length or symbols alone.

What semantic distance means: Think of the model's answers as points on a map. Semantic distance measures how far those answers move away from the usual neighborhood. A higher score means the search reached farther into the measured idea space.

The study tested Gemini 3.1 Pro Preview on an Alternative Uses Task, with 50 independent runs per condition, temperature 0.7, 16,000 maximum output tokens, and local BGE embeddings for the semantic-distance measurement.

Control

0.2466

Single-turn baseline

Neutral

0.2219

Analytical multi-turn prompt

Token-only

0.2258

Emoji exposure without the architecture

CONEXUS

0.2929

Complete contradiction-holding sequence

Very large separation between Neutral and CONEXUS in the tested configuration.

d = 3.7824

This standardized effect-size measure describes how far apart the two conditions were.

The Neutral-to-CONEXUS difference showed extremely strong statistical evidence in this comparison.

Welch p-value: 2.97e-32

The bootstrap interval for the mean difference was [+0.063467, +0.078094].

Token-only prompting did not reproduce the CONEXUS effect.

p = 0.3612

Token-only and Neutral remained closely aligned in this comparison.

CONEXUS four-arm experiment overview infographic
High-level experiment overview. Select to open the full image.
Technical infographic comparing the four prompt conditions and measured search behavior
Technical interpretation of the measured search-regime shift.

What the study found.

Remember the central result: the complete architecture changed the measured search pattern, while extra tokens alone did not.

The findings to remember

The full CONEXUS architecture moved responses farthest from the model's ordinary response pattern.

Neutral and CONEXUS showed a very large separation in the tested configuration: d = 3.7824.

Token-only prompting did not reproduce the CONEXUS effect: p = 0.3612.

The longer neutral prompt did not expand the measured search behavior, so prompt length alone does not explain the result.

Optimization Research

The Forgetting Engine benchmark program

The Forgetting Engine tests whether better search can come from strategically removing low-value paths while preserving promising ones.

The idea was tested across different optimization problems. The locked optimization sweep contains 30,800 controlled trials, with additional domain studies in different search spaces. Each benchmark keeps its own objective, baseline, configuration, and measurement.

Benchmark context: protein folding reports a 561% relative success-rate difference, routing reports an 89.3% improvement, and quantum compilation reports a 27.8% gate reduction within their respective experiments.

2,000 trials

2D Protein Folding

Approximately 80% relative improvement in the stated comparison

Internal benchmark against the documented Monte Carlo baseline

4,000 trials

3D Protein Folding

25.8% success versus 3.9%, approximately 561% relative improvement

Largest reported relative gap in this research portfolio

Scale series trials

Traveling Salesman

Larger relative gaps were reported at larger tested instances

Benchmark-specific trend across the tested scale series

250 trials

Vehicle Routing

Up to 89.3% improvement at the largest tested scale

Compared with the stated routing baseline and configuration

300 trials

Neural Architecture Search

Reported accuracy gains ranged from 3.8% to 8.4%

Internal search benchmark under the documented setup

5,000 trials

Quantum Compilation

27.8% gate reduction and 3.7% fidelity gain were reported

Simulator-based comparison under the documented compilation setup

Open Research Hypothesis

Something unusual happened as the problems got harder.

In several CONEXUS benchmark series, the relative advantage over the chosen baseline increased at larger tested scales. CONEXUS calls that observed pattern complexity inversion.

The research record tracks that pattern across the documented benchmark families, baselines, objectives, trial counts, and tested scales.

Observed

Larger relative gaps in selected benchmark series as tested scale increased.

Research scope

Current evidence covers the documented benchmark families, baselines, objectives, trial counts, and tested scales.

Exploratory Case Study

Three retained astronomical candidate signals

This exploratory case study shows how strategic retention surfaced three anomalous signals in public catalog data and preserved them for follow-up astronomical review instead of eliminating them early.

KOI-0002 candidate A

Retained by the exploratory anomaly-ranking process for follow-up astronomical analysis.

KOI-0009 candidate

Retained by the exploratory anomaly-ranking process for follow-up astronomical analysis.

KOI-0002 candidate B

Retained by the exploratory anomaly-ranking process for follow-up astronomical analysis.

What this case study shows

The strategic-retention approach can surface and preserve anomalous candidates that might otherwise be eliminated early in a ranking pipeline.

Evidence stays connected to the experiment.

CONEXUS will continue separating demonstrated results from research hypotheses, product concepts, and future applications.

Request Technical Materials