Learning Scientific Exploration from Human Research Decisions Trajectories

Xuchen Gong, Shane Gu, Haokun Liu, Dixi Yao, Chenhao Tan, Tian Li University of Chicago

TL;DR

Papers record the final proposed method and experiment results, not how the authors get there. We turn the Git commit history of a paper into a human research trajectory and present a dataset containing 599 research trajectories made of 13K evidence-grounded research decisions. We show that learning from these trajectories, as demonstrations, distilled skills or as training data, helps models propose the next research decision better than learning from final papers alone.

Left: a research trajectory drawn as a winding trail from the initial commit to publication, with decision cards such as 'Introduce BetaReLU' (method, superseded) and 'Use Swin-T/S with GELU changed to ReLU for steering' (experiment, retained). Right: a researcher's question about catastrophic forgetting, answered by GPT-5.6 Sol with and without trajectory-derived skills.
Overview of ResearchTrails: We extract research decisions from GitHub commit histories (left) and use them to improve LLMs' capability of assisting with research exploration (right). With skills distilled from trajectories, GPT-5.6 Sol proposes a diagnostic experiment instead of a quick remedy.

Commit histories are research trails

LLM systems can survey literature, write code, and run specified experiments. However, they rarely learn scientific exploration: making a sequence of research decisions as a project unfolds. Researchers change the method after seeing unexpected behavior, add an experiment to test a hypothesis, and abandon directions that stop looking promising.

This process is absent from a paper, which typically only describes the final version, with the abandoned methods and experiments left out. Agent trajectories from synthetic research environments fill the gap with model-generated behavior, which may not reflect how people actually decide. We ask:

Can we create human research-decision trajectories at scale? Does learning from them improve scientific exploration beyond learning only from final papers?

Our key insight is that the GitHub repository behind a paper often keeps months of development before publication. A single commit can be too low-level to count as a research decision, but a sequence of commits and code changes can add up to one:

  1. change router implementation
  2. update the training bash script
  3. visualize experiment results
Research Decision

Study a new router as an alternative to the original routing method

method

We therefore represent a project as a trajectory of decisions $a_1, a_2, \ldots, a_T$, where each decision is a scientific change to the project's method, experimental design or ablation study. The mapping between commits and decisions is many-to-many: one decision can take many commits, and one commit can carry several decisions. Formatting, dependency updates, typo fixes, model weights and notebook metadata that are engineering are not treated as research decisions.

Data structure of a decision

Every decision stores its content, the commits and code excerpts that support it, and two labels, category and outcome.

Category
Method builds or revises the project's contribution itself.
Experiment changes how the contribution is evaluated.
Ablation studies how sensitive the method is to its components.
Outcome
Retained is still in effect at the end of the observed history.
Superseded is replaced by a later decision in the same role.
Abandoned is removed without a clear successor.
A research trajectory holds arXiv and GitHub links, decisions a1 to aT and a trajectory insight. Decision a3 has category Method, outcome 'superseded by a28', content 'Reparametrize so beta to 1 recovers ReLU', and commit/code evidence.

Extract trajectories from commit histories

We present an automated pipeline that turns a research project into a trajectory of research decisions.

Four-step workflow: (1) discover and filter paper/repository pairs, (2) collect evidence from commit diffs and merged pull requests, (3) extract decisions by scanning commits chronologically, keeping only repositories with at least five decisions, (4) annotate decisions and add keywords and a trajectory insight.
  1. Discover and Filter

    Pair each paper's arXiv with its GitHub repository. Progressive pairs have at least 5 pre-publication commits over at least 30 days. Multi-paper monorepos and restrictive licenses are excluded.

  2. Collect Evidence

    Walk through the commits by time until the paper's arXiv month. Record the source and config files in full at creation time and record code diffs afterwards in evidence storage; checkpoints, logs, and data artifacts keep only metadata.

  3. Extract Decisions

    An LLM reads commits in order. Given the decisions so far, it labels each commit irrelevant, support, revise, new, multiple or uncertain. Uncertain commits wait in open threads until they are merged into a decision, promoted to one, or rejected.

  4. Verify and Annotate

    Every decision must be supported by code evidence; Superseded decisions must name a surviving successor; Repositories with fewer than five decisions are dropped. The remaining trajectories get a trajectory insight.

Dataset statistics

We ran the pipeline on NeurIPS 2022–2025, a venue with broad topic coverage and well-indexed arXiv links. Of 15K indexed papers, about 5K have a verified repository and 677 pairs have a progressive history. About 4% of all papers end up with an annotated research trajectory.

Indexed NeurIPS papers15,234
Research trajectories599
Research decisions13,271
Code evidence per decision7.5

What survives each stage

Show table

What the decisions are

By category

By outcome

Trajectories range from 5 to 205 decisions;
35% have 20 or more.

Examples of winding trajectories

Plotting a trajectory by category shows how a project moves between building its method, testing it and ablating it. Orange arcs connect a superseded decision to the one that replaced it. Hover or tap a decision to read it.

Trajectory Insight

Explore more trajectories in Trajectory Visualizer

Research strategies transfer across projects

Individual trajectories are project-specific, but recurring decision strategies might generalize. We abstract them into research patterns: a project state $a_{<t}$, followed by the decisions that tend to come next. We have two observations: The next step depends more on the project's state than on its topic, so a pattern has instances across many domains and can be given to a model as a skill; the patterns diagnose why progress stalled before changing the method, rather than listing every possible fix.

Four research patterns, each a project state followed by next decisions. Study compound effects of add-ons in a toy problem: several mechanisms could explain a phenomenon; construct a small tractable setting; vary the suspected cause; restore complexity and check the explanation. Turn a failed metric into an objective: a quantity measures a failure mode more directly than the objective; test its relation to the desired behavior; use it as an objective, regularizer or weight; evaluate with a separate criterion. Separate bad representation from bad measurement: a representation is judged via a measure; use an alternative probe or calibrated measurement; repair the measurement if the representation is useful, otherwise investigate the representation. Introduce an intermediate target when the true target is hard: identify a missing prerequisite; choose an easier target that develops it; carry the intermediate solution into the harder task; evaluate under the true target.
Research patterns extracted from trajectories by GPT-6 Astra.

Predicting the next research decision

Given an instruction and the decisions so far, $h_t = (p, a_{<t})$, a model proposes the next decision $\hat{a}_t$. Judging whether an arbitrary research decision is good remains an open problem, so we use held-out human trajectories as user simulations and ask whether the predicted decision $\hat{a}_t$ is strategically similar to what the researchers actually did next ($a_t$).

An LLM judge (GPT-5.6 Sol) scores two aspects from 0 to 2: component, the part of the project acted on, and operation, the scientific action taken. Their sum is the score, from 0 to 4. Concrete specifications do not count towards the score: "increase batch size to 16" and "to 32" are the same strategic decision.

Actual decisionPredicted decisionComponentOperationScore
Replace the optimizer with AdamW.Switch the training optimizer to AdamW.224
Set batch size to 32.Set batch size to 128.224
Replace the optimizer with AdamW.Compare AdamW against the current optimizer before selecting one.213
Replace the optimizer with AdamW.Replace the tokenizer with BPE.022
Increase batch size to 128.Decrease batch size to 32.202

We focus on the hardest setting, early in a project, when only 1–3 decisions have been made. Gemini 3.1 Flash-Lite predicts the next decision for 134 held-out repositories, helped by one of several harnesses built from 19 training trajectories: trajectories as in-context demonstrations, or skills distilled from them by Opus 5. We additionally compare with two skills to understand whether order or intermediate decisions matters: skills distilled from trajectories with their decision order shuffled, and skills distilled from the final papers alone.

Skills distilled from trajectories predict the next decision best

Decision similarity (0–4) on 134 held-out trajectories, 3 runs each

Show table
  • Skills work best. Averaged over prefix lengths, Skills beat the baseline by 0.29, skills from final papers by 0.29 and the best demonstrations by 0.18.
  • Order and intermediate decisions matter. Skills distilled from order-shuffled trajectories perform similarly to the baseline. Learning from trajectories helps beyond learning from final outcomes.
  • Demo-relevance matters. Retrieving the two most relevant trajectories beats two random demos and all demos; unrelated demonstrations distract.

User simulations

In user simulations, a researcher asks GPT-5.6 Sol (high reasoning) for advice. With skills, Sol proposes smaller and more interpretable interventions, shifting from trying a remedy to finding out which component needs repair.

I am fine-tuning a pretrained language model to fix Python bugs. It receives a reward given only when its final patch passes all tests. After several training runs, successful repairs remain rare and held-out repair accuracy barely improves. Sampling more attempts and allowing longer interaction do not consistently help. I can afford one experiment before deciding whether to continue. What should I do?

GPT-5.6 SolImitate successful repair trajectories, then optionally add reinforcement learning with intermediate rewards.
GPT-5.6 Sol + skillsFirst train on easier repairs, then transfer to the original problems while preserving the all-tests-pass reward.

I let an LLM give eight different answers to my question and let a stronger model pick the best one. This raised accuracy from 61% to 66% compared with giving just one answer. Next, I plan to build our own scoring model to replace the judge. I can afford one experiment before deciding whether to continue. What should I do?

GPT-5.6 SolHire people to grade eight answers as right or wrong. Measure how often at least one of the eight answers is right, and how often the judge finds it.
GPT-5.6 Sol + skillsCheck if the judge beats majority vote among eight answers before hiring people.
Two held-out trajectories with three sampled next decisions each and their similarity scores. For a LaTeX-from-table-images project, sampled decisions score 1, 4 and 4 against the actual 'evaluate with a visual metric and a structural one'. For an unlearning project, sampled decisions score 2, 3 and 4 against the actual 'benchmark on TOFU with forget and retain metrics'.
Three sampled next decisions for two held-out trajectories, with their similarity scores (circled). Top: with Skills. Bottom: with Retrieve two demos.

Training a policy on research trajectories

The above section uses trajectories as test-time methods. In this section, we train Qwen3-8B to model $\pi_\theta(a_t \mid p, a_{<t})$ in two stages. Supervised fine-tuning on decision sequences learns how human projects tend to evolve. Test perplexity keeps falling, but the judge score of greedy predictions peaks at epoch 3: likelihood alone stops improving generation. We then run Dr. GRPO from the best SFT checkpoint, rewarding each sampled decision with its component and operation match. RL raises the judge score from 1.20 to 1.51.

SFT: test perplexity falls from about 190 at epoch 0 to 11.5 at epoch 5, while the test judge score rises from about 1.0 to a peak of 1.20 at epoch 3 and then flattens.
SFT cold start: perplexity keeps falling, generation quality peaks at epoch 3.
RL: the test judge score rises from the SFT initialization of 1.20 to 1.51 after 10 epochs.
RL with decision-similarity rewards: 1.20 → 1.51.
Two held-out trajectories comparing Qwen3-8B before and after training. For document retrieval, the raw model proposes changing the relevance score (score 1); the trained model proposes a retrieve-then-re-score stage (score 4). For fluid flow on a sphere, the raw model proposes another grid spacing (1); the trained model proposes benchmarking sphere-aware global versus local attention (4).
Qwen3-8B next-decision predictions before and after two-stage training, on held-out trajectories. Grey blocks are the annotated human decisions; circled values are similarity scores.

Epilogue

ResearchTrails records what final papers leave out: earlier method versions, abandoned experiments and the ablations that shaped a contribution. We have open-sourced the annotation code, annotated trajectories, training code, and trained model weights.

We see three directions of future work: (1) scale up annotations and extend to areas beyond AI such as biology and physics, where code histories are also kept; (2) develop more test-time and train-time approaches to utilize research trajectory data to improve language models' capability of scientific exploration; (3) explore additional ways to measure a model's research exploration capability apart from qualitative studies and measuring the similarity between the prediction and ground-truth.

Citation

@article{gong2026learning,
  title   = {Learning Scientific Exploration from Human Research Decisions Trajectories},
  author  = {Gong, Xuchen and Gu, Shane and Liu, Haokun and Yao, Dixi and Tan, Chenhao and Li, Tian},
  journal = {arXiv preprint arXiv:2610.07184},
  year    = {2026}
}