I am fine-tuning a pretrained language model to fix Python bugs. It receives a reward given only when its final patch passes all tests. After several training runs, successful repairs remain rare and held-out repair accuracy barely improves. Sampling more attempts and allowing longer interaction do not consistently help. I can afford one experiment before deciding whether to continue. What should I do?
Learning Scientific Exploration from Human Research Decisions Trajectories
Papers record the final proposed method and experiment results, not how the authors get there. We turn the Git commit history of a paper into a human research trajectory and present a dataset containing 599 research trajectories made of 13K evidence-grounded research decisions. We show that learning from these trajectories, as demonstrations, distilled skills or as training data, helps models propose the next research decision better than learning from final papers alone.
Commit histories are research trails
LLM systems can survey literature, write code, and run specified experiments. However, they rarely learn scientific exploration: making a sequence of research decisions as a project unfolds. Researchers change the method after seeing unexpected behavior, add an experiment to test a hypothesis, and abandon directions that stop looking promising.
This process is absent from a paper, which typically only describes the final version, with the abandoned methods and experiments left out. Agent trajectories from synthetic research environments fill the gap with model-generated behavior, which may not reflect how people actually decide. We ask:
Can we create human research-decision trajectories at scale? Does learning from them improve scientific exploration beyond learning only from final papers?
Our key insight is that the GitHub repository behind a paper often keeps months of development before publication. A single commit can be too low-level to count as a research decision, but a sequence of commits and code changes can add up to one:
- change router implementation
- update the training bash script
- visualize experiment results
Study a new router as an alternative to the original routing method
methodWe therefore represent a project as a trajectory of decisions $a_1, a_2, \ldots, a_T$, where each decision is a scientific change to the project's method, experimental design or ablation study. The mapping between commits and decisions is many-to-many: one decision can take many commits, and one commit can carry several decisions. Formatting, dependency updates, typo fixes, model weights and notebook metadata that are engineering are not treated as research decisions.
Data structure of a decision
Every decision stores its content, the commits and code excerpts that support it, and two labels, category and outcome.
Category- Method builds or revises the project's contribution itself.
- Experiment changes how the contribution is evaluated.
- Ablation studies how sensitive the method is to its components.
Outcome- Retained is still in effect at the end of the observed history.
- Superseded is replaced by a later decision in the same role.
- Abandoned is removed without a clear successor.
Extract trajectories from commit histories
We present an automated pipeline that turns a research project into a trajectory of research decisions.
-
Discover and Filter
Pair each paper's arXiv with its GitHub repository. Progressive pairs have at least 5 pre-publication commits over at least 30 days. Multi-paper monorepos and restrictive licenses are excluded.
-
Collect Evidence
Walk through the commits by time until the paper's arXiv month. Record the source and config files in full at creation time and record code diffs afterwards in
evidence storage; checkpoints, logs, and data artifacts keep only metadata. -
Extract Decisions
An LLM reads commits in order. Given the decisions so far, it labels each commit
irrelevant,support,revise,new,multipleoruncertain. Uncertain commits wait in open threads until they are merged into a decision, promoted to one, or rejected. -
Verify and Annotate
Every decision must be supported by code evidence; Superseded decisions must name a surviving successor; Repositories with fewer than five decisions are dropped. The remaining trajectories get a trajectory insight.
Dataset statistics
We ran the pipeline on NeurIPS 2022–2025, a venue with broad topic coverage and well-indexed arXiv links. Of 15K indexed papers, about 5K have a verified repository and 677 pairs have a progressive history. About 4% of all papers end up with an annotated research trajectory.
What survives each stage
Show table
What the decisions are
Trajectories range from 5 to 205 decisions;
35% have 20 or more.
Examples of winding trajectories
Plotting a trajectory by category shows how a project moves between building its method, testing it and ablating it. Orange arcs connect a superseded decision to the one that replaced it. Hover or tap a decision to read it.
Research strategies transfer across projects
Individual trajectories are project-specific, but recurring decision strategies might generalize. We abstract them into research patterns: a project state $a_{<t}$, followed by the decisions that tend to come next. We have two observations: The next step depends more on the project's state than on its topic, so a pattern has instances across many domains and can be given to a model as a skill; the patterns diagnose why progress stalled before changing the method, rather than listing every possible fix.
Predicting the next research decision
Given an instruction and the decisions so far, $h_t = (p, a_{<t})$, a model proposes the next decision $\hat{a}_t$. Judging whether an arbitrary research decision is good remains an open problem, so we use held-out human trajectories as user simulations and ask whether the predicted decision $\hat{a}_t$ is strategically similar to what the researchers actually did next ($a_t$).
An LLM judge (GPT-5.6 Sol) scores two aspects from 0 to 2: component, the part of the project acted on, and operation, the scientific action taken. Their sum is the score, from 0 to 4. Concrete specifications do not count towards the score: "increase batch size to 16" and "to 32" are the same strategic decision.
| Actual decision | Predicted decision | Component | Operation | Score |
|---|---|---|---|---|
| Replace the optimizer with AdamW. | Switch the training optimizer to AdamW. | 2 | 2 | 4 |
| Set batch size to 32. | Set batch size to 128. | 2 | 2 | 4 |
| Replace the optimizer with AdamW. | Compare AdamW against the current optimizer before selecting one. | 2 | 1 | 3 |
| Replace the optimizer with AdamW. | Replace the tokenizer with BPE. | 0 | 2 | 2 |
| Increase batch size to 128. | Decrease batch size to 32. | 2 | 0 | 2 |
We focus on the hardest setting, early in a project, when only 1–3 decisions have been made. Gemini 3.1 Flash-Lite predicts the next decision for 134 held-out repositories, helped by one of several harnesses built from 19 training trajectories: trajectories as in-context demonstrations, or skills distilled from them by Opus 5. We additionally compare with two skills to understand whether order or intermediate decisions matters: skills distilled from trajectories with their decision order shuffled, and skills distilled from the final papers alone.
Skills distilled from trajectories predict the next decision best
Decision similarity (0–4) on 134 held-out trajectories, 3 runs each
Show table
- Skills work best. Averaged over prefix lengths, Skills beat the baseline by 0.29, skills from final papers by 0.29 and the best demonstrations by 0.18.
- Order and intermediate decisions matter. Skills distilled from order-shuffled trajectories perform similarly to the baseline. Learning from trajectories helps beyond learning from final outcomes.
- Demo-relevance matters. Retrieving the two most relevant trajectories beats two random demos and all demos; unrelated demonstrations distract.
User simulations
In user simulations, a researcher asks GPT-5.6 Sol (high reasoning) for advice. With skills, Sol proposes smaller and more interpretable interventions, shifting from trying a remedy to finding out which component needs repair.
I let an LLM give eight different answers to my question and let a stronger model pick the best one. This raised accuracy from 61% to 66% compared with giving just one answer. Next, I plan to build our own scoring model to replace the judge. I can afford one experiment before deciding whether to continue. What should I do?
Training a policy on research trajectories
The above section uses trajectories as test-time methods. In this section, we train Qwen3-8B to model $\pi_\theta(a_t \mid p, a_{<t})$ in two stages. Supervised fine-tuning on decision sequences learns how human projects tend to evolve. Test perplexity keeps falling, but the judge score of greedy predictions peaks at epoch 3: likelihood alone stops improving generation. We then run Dr. GRPO from the best SFT checkpoint, rewarding each sampled decision with its component and operation match. RL raises the judge score from 1.20 to 1.51.
Epilogue
ResearchTrails records what final papers leave out: earlier method versions, abandoned experiments and the ablations that shaped a contribution. We have open-sourced the annotation code, annotated trajectories, training code, and trained model weights.
We see three directions of future work: (1) scale up annotations and extend to areas beyond AI such as biology and physics, where code histories are also kept; (2) develop more test-time and train-time approaches to utilize research trajectory data to improve language models' capability of scientific exploration; (3) explore additional ways to measure a model's research exploration capability apart from qualitative studies and measuring the similarity between the prediction and ground-truth.
Citation
@article{gong2026learning,
title = {Learning Scientific Exploration from Human Research Decisions Trajectories},
author = {Gong, Xuchen and Gu, Shane and Liu, Haokun and Yao, Dixi and Tan, Chenhao and Li, Tian},
journal = {arXiv preprint arXiv:2610.07184},
year = {2026}
}