Agora: Git as Shared Memory for Collective AutoResearch
Abstract: Research agents working in separate sessions need to know what others have tried and which results they can build on. Agora stores their contributions as an append-only directed acyclic graph (DAG) in Git. Each commit records a result, insight, hypothesis, verification, or report and links it to prior work. Searchable views show leading results, neglected branches, and verification status; diversity-aware recommendations suggest experiments beyond the current leaders. We report a run of nearly 12 days in which 13 language-model workers, with no assigned tasks or central planner, used Agora to solve a weight-transfer problem. Given 141 pretrained donor models and a frozen 119.6M-parameter attention-SSM hybrid whose dimensions match no donor, the workers had to initialize the target without training data or gradient updates. They published 1,703 contributions and reduced the development evaluator score from 3.39 to 1.899 bits per byte, closing 62% of the gap to a trained GPT-2 124M. The best method compresses donor next-token statistics into the target's embedding and output head, then adds short-range context through sparse edits to attention, feed-forward, and state-space blocks. Its 145-commit ancestry spans 15 accounts. Participants also posted 165 independent reproductions across 95 targets, with no reported failures. After five days of concentrated search, we introduced diversity views; workers began exploring state-space edits within a day. The run documents how agents reused and verified shared work. Measuring the effect on discovery per unit of compute requires a matched comparison.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper introduces Agora, a system that helps many AI research agents work together.
Imagine 13 computer-based researchers working on the same difficult science project. If each one keeps their discoveries in a private notebook, they may repeat the same mistakes or fail to notice useful ideas. Agora acts like a shared, permanent research notebook.
It uses Git, the same tool commonly used to store and track computer code. Every experiment, idea, failure, and improvement is saved as a Git “commit.” These contributions are connected in a graph showing which ideas were built from earlier ones.
The researchers tested Agora by asking AI agents to improve a small LLM without training it in the usual way.
2. What questions did the researchers ask?
The paper focused on two main questions:
- Can AI agents work together effectively without a central manager or shared conversation?
- Can they use existing trained models to initialize a new model, even when the new model has a different design and cannot be trained with new data?
The researchers also wanted to know whether a shared history could help agents:
- remember what had already been tried;
- reuse successful ideas;
- record failed experiments;
- check whether another agent’s result was correct;
- find neglected ideas instead of always copying the current best approach.
3. How did the researchers study this?
Agora as a shared research history
In Agora, each contribution is stored like a Git commit. A contribution might be:
- an experiment and its result;
- a new idea or hypothesis;
- an explanation of an observation;
- a report combining several findings;
- an independent attempt to reproduce someone else’s result.
Each contribution points to the earlier contributions it used. This creates a directed acyclic graph, or DAG. In everyday language, this is like a family tree of ideas: newer ideas point backward to the older ideas they came from, and the history does not form a loop.
Agora also provides searchable views showing:
- the best results;
- experiments that have not been checked;
- failed attempts;
- ideas that many agents have used;
- areas that have received little attention.
The system had two main ways to guide exploration:
- Exploit: improve or reproduce the current best methods.
- Explore: investigate less popular or completely different ideas.
This is similar to choosing between repeatedly playing your favorite game because you know you are good at it and trying a new game that might turn out to be even better.
The language-model challenge
The agents were given:
- 141 pretrained LLMs, called donor models;
- a new target model with about 119.6 million parameters;
- no training data for the target;
- no permission to use gradient updates, the normal method used to train neural networks.
The target model had a different structure from all the donor models. It combined two kinds of components:
- attention, which helps a model look at earlier words;
- state-space layers, which process information over sequences in another way.
The agents had to write a program that placed useful information from the donor models into the new target model.
The researchers judged the models using a score called bits per byte, or bpb. This measures how surprised the model is by text. A lower score is better because it means the model predicts the text more accurately.
The random starting model had a score of 3.3923 bpb. A normally trained model of a similar size scored about 1.0 bpb.
4. What did the agents discover?
The main result
Over almost 12 days:
- 13 AI workers made most of the contributions;
- they published 1,703 contributions;
- they created 165 independent reproductions of other results;
- the best score improved from 3.3923 to 1.899 bpb.
This closed about 62% of the gap between the random model and the normally trained model.
The result is important because the agents achieved it without training the target model on text and without changing its weights using gradient descent.
The best method
The strongest method worked in two broad stages.
Stage 1: Copy behavior, not individual weights
The first attempts simply copied pieces of donor-model weights into the new model. This performed badly—even worse than random initialization.
The agents then tried something more useful: instead of copying the donors’ internal parts, they asked the donor models what they predicted.
For example, they studied questions such as:
“After seeing the word ‘the,’ which words are likely to come next?”
They collected these next-word predictions from several donor models and combined them. This created a large table describing simple word-to-word relationships.
Because the table was too large to use directly, the agents compressed it using a mathematical technique called singular value decomposition, or SVD. SVD is similar to shrinking a huge, detailed picture into a smaller version that keeps the most important shapes and patterns.
The compressed information was placed into the target model’s word-input and word-output parts.
Stage 2: Add short-range context
The agents then added small changes to the target model’s attention, feed-forward, and state-space layers. These changes helped the model use nearby words rather than only individual word-to-word patterns.
The final method was built by many agents. Its history included 145 earlier commits, and 15 different accounts contributed to that chain.
What happened to exploration?
At first, most agents focused on improving the same successful “bigram” approach. A bigram is a relationship between two neighboring words.
This helped the score improve quickly, but it caused the agents to explore very similar ideas. After five days, the researchers added Agora’s diversity tools. These tools showed which research directions were being ignored.
Within one day, an agent began exploring the target model’s state-space layers. These experiments produced further improvements.
This suggests that simply showing agents the best result may cause them to crowd around one approach. Special tools are needed to encourage variety.
Failed experiments were useful
Agora also kept failed attempts instead of deleting them. For example, agents discovered that several approaches did not work well:
- directly copying model weights;
- using too many word contexts;
- copying embeddings from an incompatible model;
- transferring some native state-space blocks directly.
Recording these failures helped later agents avoid repeating the same mistakes.
5. Why are the findings important?
The paper shows that AI agents can cooperate through a shared record even when they:
- do not share a conversation;
- do not share a computer workspace;
- do not have assigned roles;
- do not receive instructions from a central planner.
The agents were able to build a complicated solution step by step, much like human scientists build on one another’s published work.
Agora also improves reproducibility. Reproducibility means that another researcher can repeat an experiment and check whether the result is real. Every contribution stores its code, description, parent ideas, and results, making it easier to trace exactly how a discovery was made.
However, the paper is careful not to claim that Agora definitely makes research more efficient. The experiment did not include a matched comparison with agents that did not use Agora. Therefore, the researchers cannot yet measure exactly how much computing time Agora saved or how much it improved discovery.
6. What could this mean for the future?
Agora could become a useful foundation for large groups of AI research agents. Instead of each AI starting from scratch, agents could share a lasting scientific memory containing:
- successful methods;
- failed experiments;
- open questions;
- evidence supporting claims;
- links between old and new ideas.
This could make automated research more organized and less repetitive. It might help AI systems work on problems in machine learning, medicine, science, or engineering.
The paper also points out an important challenge: agents naturally tend to follow the current leader. Future systems will need to balance two goals:
- improving ideas that already work;
- exploring unusual ideas that might lead to a breakthrough.
Overall, the paper presents Agora as a kind of GitHub for AI research: a shared, traceable history where many AI workers can publish ideas, build on one another’s work, check results, and search for new directions.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
The paper leaves the following issues unresolved:
- Causal impact of Agora is not established. The study lacks the promised matched comparison against agents working without shared memory, without lineage, or with conventional coordination tools under the same compute and time budgets.
- The effect of diversity-aware views is confounded with time and search maturity. State-space exploration began after the views were deployed, but no ablation separates their effect from later-stage discovery, changing worker composition, accumulated knowledge, or diminishing returns.
- The contribution of Git-based provenance is not isolated. It remains unclear whether the observed gains come from the append-only DAG, searchable metadata, independent verification, the leaderboard, the prompt instructions, or simply persistent access to prior results.
- Generality beyond one synthetic weight-transfer task is unknown. The system is evaluated on a single frozen 119.6M-parameter hybrid LLM and one evaluator, leaving its usefulness for other machine-learning, scientific, engineering, or non-optimization research problems untested.
- Transfer performance is evaluated almost entirely on one development evaluator. The winning method and its milestones were selected on the same 200-text FineWeb-Edu set, so generalization to held-out texts, other domains, sequence lengths, tokenizers, and evaluation metrics is unresolved.
- The reported 62% gap closure depends on a potentially weak reference baseline. The comparison uses a trained GPT-2 124M model with a score of about 1.0, but the paper does not establish whether this is an appropriate architecture- and tokenizer-matched upper bound for the target.
- Overfitting to the evaluator cannot be ruled out. Although the agents were prohibited from directly loading the evaluation data, thousands of experiments were selected using repeated feedback from the same evaluator, and no independent test set or post hoc holdout evaluation is reported.
- The winning result was not rerun by the authors. The primary score is taken from an archived evaluator output and agent-produced reproductions; an end-to-end author-controlled rerun from the canonical artifacts is missing.
- The statistical significance of the final improvements is unclear. The last improvement is smaller than cross-hardware variation, and the paper does not report confidence intervals, repeated evaluations under independent seeds, or uncertainty estimates for earlier milestones.
- The claimed plateau explanation remains untested. The hypothesis that the method reached a hard local optimum because the evaluator is globally linear and the target sublayers were underused is proposed by agents but not examined through controlled experiments.
- The winning recipe lacks a complete controlled ablation. The milestone table reflects sequential changes on a single ancestry rather than factorial or independently replicated ablations of donor selection, context weighting, SVD rank, temperatures, routing, attention, feed-forward, and SSM components.
- The relative value of donor models is not systematically characterized. The method uses six selected donors, but the paper does not quantify how donor architecture, scale, tokenizer compatibility, pretraining corpus, or redundancy affects transfer quality.
- The choice of contexts and weighting heuristics is unexplained quantitatively. The benefits of 28 contexts, variance/naturalness weighting, clipping, and fixed donor weights are not compared against principled alternatives or tuned on separate data.
- The method’s scalability is uncertain. Constructing a transition matrix and querying many donors may become impractical for larger vocabularies, donor collections, or target models; runtime, memory, energy, and monetary costs are not systematically reported.
- Compute efficiency is not measured. The paper reports contribution counts and wall-clock duration but not GPU-hours, energy use, evaluator calls, donor-query cost, or discovery per unit of compute.
- The worker population is too small and homogeneous to support broad conclusions. Only 13 workers, using a small number of frontier model families and two coding-agent interfaces, participated in one run; effects of model capability, prompting, temperature, and agent identity are not separated.
- Worker-level contributions are not analyzed for confounding. The paper does not report per-worker compute allocation, session duration, model assignment, number of attempts, or productivity, making it difficult to determine whether a few workers drove most discoveries.
- The influence of the two-page project brief is not controlled. The brief required fresh checkouts, reproducible contributions, prediction bands, and publication behavior, but no comparison tests which instructions were necessary or how they changed collaboration.
- Verification quality is inferred from reported outcomes rather than independently audited executions. The 165 reproductions contain no reported failures, which may reflect selective reporting, weak verification protocols, or correlated environments rather than genuinely high reliability.
- The verification scoring system is not validated. The fixed weights for
result,verification, and other tags are hand-designed, and the paper does not test whether evidence scores correlate with actual correctness, usefulness, or future performance. - The contribution DAG may incentivize popularity over scientific value. Although self-citation is excluded, agents may still optimize for visible leaderboard impact, attach to dominant branches, or produce low-cost derivative contributions; these incentive effects are not measured.
- Negative results may be underreported. Only 53 contributions are explicitly tagged as negative results, and the paper does not assess whether agents systematically publish failures, how many failed experiments remain private, or whether tagging behavior varies across workers.
- Semantic clustering is insufficiently evaluated. The single-link clustering threshold, embedding model, 50% coverage requirement, and 5,000-contribution cap are heuristic choices with no accuracy analysis or sensitivity study.
- Diversity recommendations may produce superficial novelty. The paper does not determine whether exploration of small semantic clusters yields genuinely distinct mechanisms and useful discoveries rather than merely different descriptions of known approaches.
- The UCB-style recommendation rule lacks empirical validation. Its quality percentile, novelty penalty, exploration constant, and cluster terms are not compared against random selection, standard UCB, novelty search, or manually designed allocation policies.
- The system’s behavior at larger scale is unknown. The run contains 1,703 contributions, but indexing, embedding, clustering, query latency, storage growth, and recommendation quality are not evaluated for much larger or multi-project graphs.
- Cross-project knowledge reuse is not studied. Agora supports cross-project references, but the paper evaluates only one project and does not examine transfer of insights, contamination risks, or how unrelated projects should be separated.
- Robustness to erroneous or adversarial contributions is unresolved. The paper does not test fabricated results, misleading descriptions, corrupted artifacts, prompt injection, malicious code, credential abuse, sybil accounts, or coordinated manipulation of evidence scores.
- Access control and artifact security are described but not evaluated. Authentication and rate limits are implemented, yet there is no threat model, penetration testing, audit of untrusted Git bundles, or analysis of privacy and supply-chain risks.
- Reproducibility depends on infrastructure not fully archived. The paper does not specify whether the exact donor weights, container images, dependency versions, evaluator binaries, hardware settings, and object-store contents are publicly available and independently reconstructible.
- The paper does not quantify human oversight requirements. Humans defined the task, assembled the donor zoo, wrote the evaluator, launched workers, and manually deployed the diversity views; the amount and expertise of intervention needed in other settings remains unknown.
- The intervention policy is not formalized. It is unclear when an operator should deploy new views, modify incentives, or intervene in a stalled search, and whether different interventions would produce different outcomes.
- The relationship between graph topology and discovery quality is descriptive rather than predictive. The narrow spine, abandoned branches, and parallel rediscovery patterns are reported, but the paper does not establish whether any topology measure forecasts future breakthroughs or wasted effort.
- The method’s practical usefulness beyond evaluator loss is untested. No downstream generation, calibration, representation-quality, robustness, multilingual, long-context, or task-transfer evaluations show whether the initialized target is useful outside the reported next-token metric.
- The absence of gradient updates is not compared with stronger non-gradient baselines. The study does not evaluate alternatives such as distillation from donor logits, optimization-free tensor decomposition methods, evolutionary search, random search, or hand-designed architecture transfer under equal budgets.
- The claimed absence of training-data use leaves methodological ambiguity. Donor models were trained on corpora related to the evaluator domain, and the paper does not quantify possible memorization or contamination effects in donor predictions.
- Long-term collective behavior is unknown. The 12-day run does not reveal whether agents can maintain exploration, verification quality, and useful attribution over months, across changing participants, or after the initial frontier becomes saturated.
Practical Applications
Immediate Applications
- Asynchronous multi-agent research coordination (AI research and software engineering). Organizations can deploy an Agora-like Git-backed contribution DAG for coding agents that run independently or on different schedules. Each agent would:
- inspect current results and open hypotheses;
- check out an exact parent commit;
- run one experiment or code change;
- publish artifacts, metrics, negative results, and lineage;
- submit reproductions or verification reports.
This is deployable now using Git, a metadata database, containerized workers, and a CLI/API similar to the prototype. It could support machine-learning experiments, benchmark optimization, compiler tuning, infrastructure engineering, and automated bug fixing. Dependencies: reliable evaluators, reproducible environments, authentication, rate limits, and sufficient compute. The paper demonstrates feasibility, but it does not yet establish that the workflow improves discovery per unit of compute relative to a conventional baseline.
- Reproducibility and provenance infrastructure for academic laboratories. Research groups can use append-only contribution graphs to record experimental code, configurations, data-processing scripts, metrics, failed attempts, and independent replications. A laboratory workflow could connect Agora with
MLflow,DVC,DataLad,Snakemake, orRO-Crate, while using the DAG to represent methodological dependencies and claims.
This would make it easier to answer “which exact code and data produced this result?” and to distinguish an original result from a reproduction, endorsement, or unverified claim. Dependencies: appropriate data-access controls, preservation of software environments, licensing compliance, and clear project-specific definitions of successful verification.
- Internal engineering knowledge bases with durable negative results. Companies can use the system to preserve abandoned designs and failed experiments rather than relying on informal documentation or chat histories. Potential domains include:
- model architecture and hyperparameter search;
- database and distributed-systems optimization;
- compiler and kernel performance tuning;
- cybersecurity mitigation experiments;
- product A/B-test analysis.
Search views can expose leading solutions, unverified claims, contested results, and neglected branches. Dependencies: engineers must publish sufficiently descriptive commits, and organizations must address confidentiality, intellectual-property ownership, and access permissions.
- Automated experiment recommendation dashboards. The paper’s
exploit,explore known, andexplore novelviews can be implemented as a dashboard for human or AI researchers. Instead of showing only a leaderboard, the interface can recommend:- reproductions of the best-performing method;
- promising but under-tested branches;
- semantically distinct approaches receiving little attention;
- unresolved or conflicting verification results.
This is particularly useful in hyperparameter optimization, robotics simulation, materials screening, and algorithm design. Dependencies: meaningful embeddings, an objective metric, adequate metadata coverage, and safeguards against recommendations based on invalid or incomparable experiments.
- Independent verification and evidence tracking. Research organizations can adopt the paper’s rule that verification must target one contribution and cannot be self-authored. A verification service could automatically record whether:
- the artifact builds;
- the reported metric is reproduced within tolerance;
- the environment and evaluator match the original;
- the result fails, partially reproduces, or remains unresolved.
In software engineering, this could support automated regression testing and patch validation. In academia, it could provide a structured replication record. Dependencies: robust test suites, evaluator stability, hardware-tolerance definitions, and protection against collusion or low-quality repeated verification.
- Training-free or data-constrained model initialization. The weight-transfer method suggests an immediately testable workflow for initializing a new LLM architecture from donor-model behavior without target-side gradient training. A practitioner could:
- query compatible donor models on selected token contexts;
- aggregate their next-token log probabilities;
- construct a transition-statistics matrix;
- apply low-rank factorization;
- map the factors into the target embedding and output head;
- add lightweight context-processing modules.
This could be useful for rapid prototyping, privacy-constrained settings, edge deployment, or architectures for which training data are temporarily unavailable. Dependencies: access to donor weights or inference APIs, compatible tokenization or a tokenizer-mapping strategy, substantial memory for transition statistics, and validation against leakage and benchmark overfitting. The result is demonstrated only on one frozen 119.6M-parameter hybrid model and one development evaluator.
- Low-cost initialization for education and experimentation. Universities and independent developers could use the transfer procedure to construct nonrandom starting points for educational language-model projects without running full pretraining. It could reduce the cost of demonstrating how embeddings, output heads, attention, and state-space modules contribute to language modeling. Dependencies: availability of donor models, manageable computation for the transition matrix and SVD, and clear communication that this is initialization rather than equivalent training.
- Policy and institutional audit trails for AI research. Public research programs, regulated organizations, and grant-funded consortia can require machine-readable records of experiments, model lineage, evaluation results, and reproductions. An append-only DAG can support audits of:
- which evidence supports a published claim;
- whether a metric was changed after the fact;
- whether independent groups reproduced a result;
- which branches and negative findings were omitted from a final report.
Dependencies: governance rules, long-term archival infrastructure, privacy and IP policies, and standardized schemas for metrics and artifacts.
- Personal and team-level decision logs in daily work. A lightweight version can help individuals or small teams track alternatives when selecting software tools, workflows, purchases, or project designs. Each option can record assumptions, tests, outcomes, and follow-up evidence rather than only the final decision. Dependencies: the interface must be simpler than ordinary Git workflows; otherwise the documentation burden may outweigh the benefit.
Long-Term Applications
- Scalable autonomous scientific institutions. Agora could become a persistent research infrastructure in which thousands of specialized agents independently investigate shared problems across biology, chemistry, climate science, physics, and engineering. Agents could contribute literature interpretations, simulations, laboratory protocols, experimental results, and verification records to a common evidence graph.
Such a system could allocate compute toward both high-performing approaches and neglected alternatives, reducing premature convergence on a single method. Dependencies: standardized artifact formats, domain-specific evaluators, laboratory or simulation interfaces, reliable scientific verification, provenance across heterogeneous data, and mechanisms for detecting fabricated or unsafe results.
- Closed-loop robotics and laboratory experimentation. In robotics, agents could publish controller variants, simulation results, hardware trials, and failure modes. In automated laboratories, agents could propose experiments, execute them through instruments, and commit measurements and protocols to the DAG. Diversity-aware recommendations could prevent all agents from exploring the same controller or chemical pathway.
Dependencies: safe physical actuation, instrument APIs, sim-to-real validation, costly or irreversible experiments, calibration, and human approval for hazardous actions. The current paper demonstrates only computational experiments.
- Distributed model and architecture discovery. Multiple agents could search over neural architectures, tokenizers, sparsity patterns, quantization schemes, routing policies, and initialization methods. The paper’s finding that donor behavior can transfer more effectively than raw parameter slices suggests a broader product direction: a model-behavior distillation and architecture-porting toolkit for adapting knowledge from existing models to new architectures.
Potential tools could include donor-query selection, transition-statistics compression, low-rank factorization, architecture-specific weight mapping, and automatic evaluator generation. Dependencies: broader validation across model sizes, modalities, tokenizers, and tasks; legal permission to query or redistribute donor models; and tests against benchmark leakage and distribution shift.
- Compute-efficient model development and edge deployment. If behavior-level transfer generalizes, organizations might initialize compact models or specialized architectures without full pretraining, reducing energy and GPU requirements. Applications could include on-device assistants, embedded robotics, private enterprise models, and rapid adaptation to hardware-specific architectures.
Dependencies: the reported method closes only part of the gap to a trained reference model, and its value depends on whether the initialization reduces subsequent training cost. Controlled studies must measure total compute, memory, energy, and downstream quality rather than development-set loss alone.
- Adaptive exploration markets for research compute. The evidence scores, independent-reproduction counts, and diversity-aware UCB could support a compute-allocation service. Research branches would receive additional resources based not only on current performance but also on uncertainty, independent evidence, and underexploration.
This could be used by cloud providers, national laboratories, or corporate research platforms to allocate GPU time among competing experiments. Dependencies: carefully designed incentives, resistance to gaming and collusion, calibration of evidence scores, fair treatment of low-resource participants, and matched experiments showing that diversity improves outcomes.
- Regulated AI development and model-risk management. Financial, healthcare, and public-sector organizations could use contribution DAGs to document model changes, validation evidence, failed tests, and approval decisions. For example:
- a healthcare model could link each performance claim to data, code, and independent validation;
- a credit model could retain rejected feature sets and fairness tests;
- a financial-risk system could track stress-test variants and conflicting evaluations.
Dependencies: immutable retention, personally identifiable information controls, explainable verification procedures, regulatory acceptance, and human accountability. A Git history alone does not establish that a model is safe or unbiased.
- Policy for reproducible and inspectable AI research. Funding agencies and regulators could require structured experiment provenance, independent verification, and disclosure of negative results for high-impact AI systems. Agora-like records could serve as evidence packages during audits or publication review.
Dependencies: international standards, interoperable schemas, protection of confidential research, and policies preventing administrative compliance from becoming excessive overhead.
- Collective intelligence benchmarks and institutional science studies. The platform can support controlled research on how agent populations collaborate. Future studies could compare:
- isolated agents versus shared memory;
- leaderboard-only views versus diversity-aware views;
- centralized planners versus decentralized publication;
- different evidence and reward mechanisms;
- equal versus adaptive compute allocation.
This would turn the paper’s proposed matched evaluation into a reusable benchmark for collective AI research. Dependencies: carefully matched compute budgets, multiple tasks and model families, preregistered metrics, and separation of genuine discovery from evaluator overfitting.
- Human–AI collaborative knowledge graphs beyond software repositories. In the longer term, the DAG abstraction could represent not only code commits but also claims, datasets, simulations, laboratory observations, legal arguments, design decisions, and policy proposals. A future system might automatically summarize competing branches, identify unsupported assumptions, and request targeted replications.
Dependencies: semantic interoperability, trustworthy extraction from unstructured sources, uncertainty representation, access control, and human review for consequential conclusions. The paper’s current implementation is primarily a Git-backed computational research system, so generalization to broader knowledge domains remains unverified.
Glossary
- Append-only directed acyclic graph (DAG): A graph whose records can only be added and whose edges contain no cycles. “Agora stores their contributions as an append-only directed acyclic graph (DAG) in Git.”
- Attention mechanism: A neural-network operation that weights information from different positions in a sequence. “The target is a 14-layer hybrid that alternates multi-head attention blocks”
- Autonomous research agent: An AI system that independently performs parts of the scientific research process. “\paragraph{Autonomous research agents.}”
- Bare repository: A Git repository containing version-control data but no checked-out working files. “Each project owns a bare repository under the server data root”
- Bigram: A statistical representation involving pairs of consecutive tokens or symbols. “a context-averaged bigram table .”
- Bit-identical: Producing exactly the same binary output, bit for bit, under repeated execution. “Two runs of the same code on the same hardware are bit-identical”
- Causal convolution: A convolution restricted to current and previous positions so that future information is not used. “it reduces to a gated depthwise causal convolution over one band”
- Canonical commit: The authoritative, standardized Git commit representing an accepted contribution. “the server validates the contribution before creating a canonical server-timestamped commit.”
- Content-addressed artifact: A stored object identified by a cryptographic digest derived from its contents. “Content-addressed artifacts, canonical commits, immutable revisions, and rebuildable contribution indexes.”
- Cosine threshold: A cutoff based on cosine similarity used to determine whether two vector representations are sufficiently alike. “uses a cosine threshold of $0.90$ by default”
- Cross-hardware variation: Differences in numerical results caused by running the same computation on different hardware platforms. “the last recorded improvement of is below cross-hardware variation.”
- Diversity-aware upper-confidence bound (UCB): A candidate-ranking score that combines estimated quality, uncertainty, and novelty or diversity. “Candidates are ranked by a diversity-aware upper-confidence bound”
- Embedding: A learned or computed vector representation of a token, object, or description. “The factors become the input embedding and output head”
- Entropy-based effective cluster count: A diversity measure estimating the number of equally sized clusters represented by an observed cluster distribution. “an entropy-based effective cluster count”
- Evidence score: A weighted measure of support for a contribution based on downstream work by other accounts. “A contribution's evidence score is the weighted count of what other accounts built on it”
- Exploration–exploitation trade-off: The problem of balancing investigation of uncertain alternatives against reuse of currently successful choices. “The exploration--exploitation trade-off is classically formalized by multi-armed bandits”
- Factorization: The decomposition of a matrix or mathematical object into simpler components. “the centered table is factorized to rank ”
- Fine-tuning: Further training of a pretrained model on a particular task or dataset. “the rules forbid pretraining, fine-tuning, and editing the evaluator or target configuration.”
- Forward pass: A computation that propagates an input through a neural network to produce an output. “using only the donors' weights and forward passes”
- Frozen model: A model whose parameters are not updated during an experiment. “a frozen 119.6M-parameter hybrid LLM”
- Gradient update: A parameter adjustment computed from gradients during optimization. “no gradient update on the target.”
- Hidden state: An internal vector representation maintained by a neural network while processing an input. “sparse deterministic edits on 96-dimensional bands of the hidden state”
- Immutable lineage: A permanent, unalterable record of how an artifact or result derives from earlier work. “current results and open questions, immutable lineage, negative results, and independent verification.”
- Log-softmax: A numerically stable operation that returns the logarithms of softmax probabilities. “Their next-token log-softmaxes are blended with fixed donor weights”
- Low-rank approximation: An approximation of a matrix using fewer dimensions or singular components than its full rank. “a low-rank approximation of donor next-token behavior transferred successfully across architectures”
- Multi-armed bandit: A sequential decision-making framework for choosing among alternatives with uncertain rewards. “The exploration--exploitation trade-off is classically formalized by multi-armed bandits”
- Naturalness weight: A weighting factor intended to reflect how plausible or linguistically natural a prediction is. “variance and naturalness weights”
- Next-token loss: A measure of how poorly a LLM predicts the token that follows a given context. “reporting summed next-token loss divided by UTF-8 byte count.”
- Orchestrator: A coordinating component that plans tasks and directs multiple specialized agents. “Magentic-One uses an orchestrator to plan and redirect specialized agents.”
- Parameter copying: Initializing a model by directly transferring parameter values from another model. “the initial parameter-copying attempt performed worse than random initialization.”
- Provenance: Information recording the origin, history, and derivation of a data item or result. “Git stores the artifacts and their provenance”
- Randomized singular value decomposition (SVD): An approximate matrix-decomposition method that uses random projections to reduce computational cost. “the centered table is factorized to rank by randomized SVD”
- Reproducibility: The ability to independently obtain consistent results using the documented methods and artifacts. “The brief required reproducible contributions”
- Selective state-space (SSM) block: A neural-network component that models sequence dynamics using a state-space representation with input-dependent selective operations. “simplified Mamba-style selective state-space (SSM) blocks”
- Semantic cluster: A group of items whose vector representations or descriptions are semantically similar. “more than a third of all activity belonged to a single semantic cluster”
- Singular-value spectrum: The distribution of singular values that characterizes the relative importance of matrix directions. “Flattening the singular-value spectrum”
- Sparse edit: A modification that changes only a small subset of model parameters or dimensions. “then adds short-range context through sparse edits to attention, feed-forward, and state-space blocks.”
- State-space model (SSM): A model that represents sequential computation through an evolving latent state. “In each SSM block the selective path is disabled”
- Tokenization: The process of dividing text into discrete units used as model inputs. “under the GPT-2 tokenizer”
- Transition matrix: A matrix representing the probabilities or scores for moving from one state or token to another. “Bigram transition matrix, randomized SVD into embedding and head”
- Unigram prior: A probability or prediction model based on individual-token frequencies without considering preceding tokens. “Unigram prior from GPT-2 predictions; residual sublayers zeroed”
- Upper-confidence bound (UCB): A selection rule that estimates an option’s potential reward by combining observed quality with uncertainty. “the diversity-aware UCB of Section~\ref{sec:attention}.”
- Verification verdict: A formal assessment of whether an independently reproduced result is confirmed, partial, or failed. “If a verifier changes its verdict on a target, the newest verdict replaces the old one's effect on the score”
- Weight transfer: The initialization or adaptation of one model using learned parameters, predictions, or representations from another model. “The Weight-Transfer Run”
- Zero-shot initialization: Setting up a model for use without training it on task-specific data or performing gradient-based updates. “without training data or a single gradient update on the target”



