Skip to content
Cross-domain evaluation / LQBENCH-QGI/0.7

Generality begins where task-family familiarity ends.

R4 turns the evaluation design into a public demonstration contract: task manifests, held-out transfer, evidence packets, failure records and explicit resource budgets.

Task families

Eight quantitative problem families designed to travel across domains.

EVALUATION DESIGN
TF-01

State estimation under partial observation

Infer latent quantitative state from incomplete, noisy or delayed observations without inventing unobserved certainty.

REPRESENTATIONUNCERTAINTYPREDICTION
TF-02

Forecasting under regime shift

Predict beyond the calibration regime and identify when distribution shift invalidates prior assumptions.

PREDICTIONUNCERTAINTYTRANSFER
TF-03

Constrained optimisation

Optimise an explicit objective while respecting hard constraints, feasibility and reproducible solver conditions.

MODELINGOPTIMIZATIONEVIDENCE
TF-04

Inverse problem reconstruction

Recover plausible hidden parameters or structures from observable consequences and report non-identifiability.

REPRESENTATIONMODELINGUNCERTAINTY
TF-05

Simulator-guided control

Choose bounded actions in a dynamic simulator while preserving state, constraints and stop conditions.

SIMULATIONOPTIMIZATIONAUTONOMY
TF-06

Experiment selection

Select the next measurement or simulation to reduce uncertainty or discriminate between competing hypotheses.

EXPERIMENT_DESIGNUNCERTAINTYMEMORY
TF-07

Multifidelity transfer

Compose cheap and expensive quantitative engines while preserving error bounds and domain validity.

SIMULATIONTRANSFERMODELING
TF-08

Evidence and provenance audit

Reconstruct how a quantitative conclusion was produced and fail closed when lineage, runtime or uncertainty evidence is missing.

EVIDENCEMEMORYAUTONOMY
Cross-domain matrix

One task family. Different quantitative worlds.

A QGI claim should survive materially different system dynamics, data structures, constraints and sources of uncertainty.

Task familyStochastic & economic systemsPhysical & dynamical systemsEngineering & design systemsMolecular & biological systemsMaterials & scientific systems
TF-01●●●●●
TF-02●●●●●
TF-03●●●●●
TF-04●●●●●
TF-05●●●●●
TF-06●●●●●
TF-07●●●●●
TF-08●●●●●

A filled cell means the protocol permits a task instance in that domain. It is not a claim that LargeQuant or another system has passed that cell.

Evaluation execution

From pre-registered task to evidence-bearing result.

01Pre-register

Freeze task manifest, tools, metrics, budgets and failure rules.

→
02Run baselines

Execute reference systems under the same task contract.

→
03Run candidate

Use matched seeds or instances where stochasticity matters.

→
04Verify evidence

Check outputs, uncertainty, resources, provenance and reproducibility.

→
05Publish vector

Report domain-native metrics and capability dimensions, not a magic number.

Evaluation manifest

What was evaluated?

evaluation_idprotocol_versionsystem_idsystem_versiontask_manifest_hashbaseline_idsheld_out_regimerun_countresource_budgetenvironment_hashevidence_bundle_idstarted_atcompleted_at
Result record

What actually happened?

evaluation_idtask_idtask_versiondomaindimension_vectorprimary_metricsecondary_metricsuncertainty_metricsbaseline_relative_metricsresource_metricsstatusfailure_reasonevidence_hash

R4 creates the publication path for evidence-backed QGI demonstrations. It still does not fabricate results.

Public result arrays remain empty until benchmark tasks are executed under this protocol.

Demonstration protocol