Skip to content

Limitations and Honest Assessment

What the repository currently establishes, what it proposes, and what would change our mind.


Current character of the project

Systems & Intelligence is a living research space. Its current foundation is a compositional probability model of typed stochastic processes. That basis is established mathematics: standard Borel spaces, Markov kernels, composition, probability, information theory, and statistical decision theory.

The repository's possible contribution is not a new universal mathematics of intelligence. It is the organization of several bounded questions:

  1. Which candidate process models remain identifiable from declared observations and interventions?
  2. Which constraint architectures keep capable, coupled systems viable and correctable?
  3. Under what conditions can heterogeneous participants produce solutions that none could reach alone?
  4. How can repeated practices stabilize action, identity, and culture without treating those words as context-free mathematical objects?

The first two have executable toy models. The latter two are currently conceptual research directions.

What is on firm ground

  • The reconstructed process foundation is an instance of established categorical probability, not new mathematics.
  • An observed trace can fail to identify a unique latent process. The foundation gives a simple hidden-extension construction, and the inverse-reconstruction benchmark exhibits finite equivalence classes in selected model families.
  • Observation, conditioning, and causal intervention are distinct once their interfaces are declared.
  • Identity, learning, and intelligence require additional tests, tasks, losses, resources, or intervention families; they do not follow from dynamics alone.
  • Phenomenal consciousness is not derived from functional organization in this repository.
  • The numerical outputs reported for the current benchmark and agentic toy experiments are reproducible from the checked-in code.

These statements are modest. Most are consequences or demonstrations of known theory.

What is conditional on a model

The Viable Corridor couples replicator, Kuramoto, regulation, and substrate equations. Its necessity results hold under the assumptions named in the paper. Its capability-loading pattern appears in two synthetic implementations. This is evidence about those models, not yet about production AI systems, organizations, or civilizations.

In particular:

  • \(\lambda_2>0\) establishes connectivity of an undirected graph; it does not by itself establish robustness to node loss.
  • \(K>K_c\) has meaning only for a specified oscillator model, topology, frequency distribution, and limiting regime.
  • a TEO dissipation proxy is not a direct measurement of ecological or computational thermodynamics;
  • the conjunction defining the corridor is not yet proved sufficient;
  • calling the selected constraint bundle love is normative framing, not a theorem.

What the current experiments do not establish

Inverse reconstruction

The benchmark measures search and identifiability inside small declared languages and model families. Exponential enumeration curves are properties of those encodings and algorithms. They do not prove a general lower bound and do not settle P versus NP.

Interventions reduce a consistent-model class when the true process and a discriminating query lie inside the declared setup. This does not mean intervention always identifies a unique real mechanism.

Agentic identity suite

The suite distinguishes hand-built architectures under selected perturbations. Identity Persistence and Δ-Kohärenz are instruments for those tests, not universal measures of identity or consciousness.

Experiment 8 omits a constant bias from both compared filters. Its failure demonstrates model omission in that setup. It does not prove that a single observation channel makes the bias structurally unidentifiable. A proper identifiability study must vary initial-state knowledge, priors, reference signals, and an augmented state model.

Preference consistency

The utility-engineering scripts currently feed hard-coded choices into a graph diagnostic. They do not query live models. An intransitive response graph can reflect framing, sampling, context, or aggregation; it does not uniquely reveal an internal utility function, irrationality, reward hacking, or safety risk.

The older ChatGPT-versus-Claude note records an unreproducible manual exploration. It is research history, not empirical evidence.

Claims deliberately not made

This repository does not currently establish:

  • a universal scalar measure of intelligence;
  • that an unqualified “generator” is a mathematical primitive;
  • that recovering a process is generally harder than running it;
  • that emergence has one scale-invariant equation;
  • that societies and AI systems are mathematically identical;
  • that any current AI system is conscious;
  • that Gödel's incompleteness theorem requires human oversight or prevents self-modeling;
  • that thermodynamics alone determines ethical values;
  • that a toy simulation predicts a deployment outcome.

When an older exploratory page sounds stronger, this assessment and What This Project Does NOT Claim govern.

The main methodological risk

The repository is unusually rich in cross-domain mappings. That makes it generative, but also makes analogy easy to mistake for identity. A mapping earns stronger status only when it declares:

  1. source and target objects;
  2. preserved relations;
  3. observation and intervention procedures;
  4. estimated parameters and uncertainty;
  5. a prediction or failure condition not guaranteed by construction.

Shared vocabulary is not enough.

What must come next

  1. Model-identification controls: compare search methods under matched languages, compute budgets, noise, and out-of-family cases.
  2. Experiment 8 controls: add augmented-state estimators and vary what is known about initial state, bias, and reference signals.
  3. Real-model preference study: preregister prompts and exclusions; repeat across orderings, temperatures, model versions, and sessions; publish raw responses.
  4. External viability tests: calibrate variables to one bounded real system before making civilizational claims.
  5. Cooperative-intelligence experiment: compare individual and mixed teams under equal budgets, real revision authority, and independent verification.
  6. Culture bridge: operationalize recurrence, transmission, correction, and power; test whether stable practices predict future action better than stated knowledge alone.
  7. External review: seek specialists in dynamical systems, causal inference, information theory, cognitive science, anthropology, and organizational research.

Falsification posture

The project should shrink when evidence requires it.

  • If passive data identifies the relevant process as reliably as intervention under matched access, the intervention claim weakens.
  • If capability does not load multiple constraints in calibrated systems, the corridor's broader architecture claim weakens.
  • If identity instruments fail to predict held-out behavior better than simpler baselines, retire them.
  • If heterogeneous cooperation adds no reachable solutions after coordination cost, the separatrix hypothesis fails.
  • If recurring practices add no predictive value beyond knowledge, incentives, and resources, the proposed culture bridge fails.

The repository is successful if it makes these failures visible, not if every earlier idea survives.

References