Illustration: an AI agent with a flame of ideas brainstorming with a group of researchers at night

AI Night-Scientist

Reinforcing Agentic Creativity inScientific Ideation with Night Science

1University of Illinois Urbana-Champaign 2Microsoft 3Microsoft Research *Corresponding: pk36@illinois.edu, sjauhar@microsoft.com

Why night science?

LLMs are hyper-focused on day science. Discovery also needs night science.

LLMs tend to give the most likely answer. That works for math and code, but it makes their research ideas predictable and similar to each other.

Day science

Careful and step by step. Form a hypothesis, test it, refine it.

Night science

Loose and exploratory. Borrow ideas from distant fields, spontaneous debates, follow hunches, question what everyone assumes. Breakthroughs like chemotherapy came from a moment like this.

AI Night-Scientist uses reinforcement learning to teach a model when and how to step away from the obvious answer.

+27.8%
more kinds of research ideas
paradigm diversity, relative to the base model
+14.9%
more kinds of contributions
contribution-type diversity, relative to the base model
+32.0 pts
predicted citation impact
win rate, relative to the base model
+66.2 pts
originality
win rate, relative to the base model
Abstract ShowHide

Large language models (LLMs) excel at structured, verifiable tasks, but their low-entropy bias can produce homogeneous and predictable outputs, limiting their utility for open-ended scientific ideation. Effective discovery, however, spans a broader creative spectrum: from structured day science to loosely structured, serendipitous night science that reaches ideas beyond those typically considered. We introduce AI Night-Scientist, an agentic framework that uses reinforcement learning to teach models when and how to depart from predictable reasoning. Grounded in cognitive science, we model creativity along three axes: action (what to do and how creatively), process (when to explore versus exploit), and outcome (the novelty and usefulness of the resulting idea). We use these axes to train models with GRPO, exposing them to varying degrees and forms of creativity throughout training. This produces substantially more diverse scientific proposals, expanding the range of research directions by 27.8% and contribution types by 14.9% over the base model. It also improves predicted citation impact by up to 32.0 percentage points and originality by 66.2 points. These gains cannot be reproduced by simply increasing decoding temperature; instead, we find that semantic guidance specifying what kind of creativity to pursue is critical. Overall, our results suggest that creativity is a learnable, multi-level ability that can be shaped to help researchers reach ideas beyond those typically explored by LLMs.

The idea

Creativity happens at three levels

Most systems only judge whether the final idea is creative. We also let the model decide how creative to be at each step, and when.

HOW

Action-level

How creative is each step? A literature search can stay close to the topic or reach into a distant field.

WHEN

Process-level

When should the model take a creative leap, and when should it build on what it already has?

WHAT

Outcome-level

Is the final idea new, while still relevant and doable?

Example for a research problem on education with LLMs: action-level creativity compares a low-creativity and a high-creativity search; process-level creativity shows a sequence of search, write, debate, spark and write steps; outcome-level creativity scores the final idea on relevance, novelty and feasibility.
An example for the problem “advancing education using LLMs.” Each level can be dialed from low to high creativity.

How it works

Given a research problem, the model writes a research proposal in up to five steps.

  1. 1
    Choose. Pick an action and a creativity level, like “search, level 4.”
  2. 2
    Act. Search arXiv, simulate a debate, or challenge an assumption. Add to the agent memory.
  3. 3
    Repeat, then write. Keep going until the model stops or hits max steps, then hand in the proposal.
  4. 4
    Learn (training only). Score the proposal. RL (GRPO) makes the choices behind good proposals more likely.
Framework diagram: the agent selects an action and creativity level from an action pool, or gets a random swap with probability Pr_swap; it executes the action, such as an arXiv search, updates its context history, and repeats until it chooses to complete; the final proposal is split into atomic ideas and scored for precedence, feasibility and relevance.
The full loop. Sometimes the chosen action is swapped for a random one (the dice), so the model tries things it would not pick on its own.

Five actions, each defined with creativity levels

Search, Debate and Spark have levels from 1 to 5 (levels 1, 3 and 5 shown). Write and Stop have no creativity levels, so the final proposal reflects the exploring that came before it.

ActionWhat it doesLevel 1 → 3 → 5
aSearch Write a query and retrieve arXiv papers. 1: the proposal’s own topic
3: nearby topics
5: distant fields and bigger questions
bDebate Pick who is in the room and what to discuss, then simulate the conversation. 1: a colleague in the same area
3: someone in the same field, different topic
5: experts from distant fields, open-ended
cSpark Find an assumption (Bit), flip it (Flip), and turn the flip into a new idea (Spark). 1: a small assumption in the proposal
3: an assumption behind the approach
5: a big assumption the whole field makes
dWrite Turn everything so far into a proposal draft. Fixed
eStop Finish and hand in the proposal. Fixed

How proposals are scored

We split each proposal into its individual ideas and grade each one, checking against related papers.

Precedence

Is it different from existing work?

Feasibility

Is the plan specific and believable? We keep the weakest idea's score, since one bad step can sink a project.

Relevance

Does it address the original problem?

Outcome reward  Rout = ⅓ (Precedence + Feasibility + Relevance)

We also tried rewarding the steps themselves, scoring each one on whether it explored something new and whether it fed into the final proposal (Rpo). Our main models use the outcome reward only; the results below compare the two.

Results

Better and more varied proposals

We train Qwen3-8B and Qwen3-14B on 4,414 NSF grant titles in computer science, engineering and math, and test on 491 more. The model only sees the title.

(a) Proposal quality

MethodCitation %Originality %
1Zero-shot
Llama-3.1-8B1.611.15
Qwen3-8B0.462.53
Qwen3-14B1.152.53
GPT-4.13.9012.18
Temperature
Qwen3-8B + Temp2.070.69
GPT-4.1 + Temp2.999.89
ReAct
Qwen3-8B + ReAct2.336.54
GPT-4.1 + ReAct3.4510.11
Creativity levels, no RL
Llama-3.1-8B + Creative1.106.52
Qwen3-8B + Creative1.894.25
Qwen3-14B + Creative0.714.76
GPT-4.1 + Creative3.0014.02
With RL
GIANTS-4B1.6235.57
2Qwen3-8B + Temp11.5219.82
3Qwen3-8B + ReAct24.9446.60
3Qwen3-8B + Search + Write24.4733.33
4Ours
AI Night-Scientist-8B29.8956.32
AI Night-Scientist-14B33.1868.68

Best in bold, second-best underlined.

How to read it. Each proposal goes head to head with a reference proposal rebuilt from the real, funded NSF project. The numbers are how often ours wins. Citation: SciJudge predicts which would be cited more. Originality: GPT-5.1 picks the more original one, given the closest existing paper.
1

Without training, models almost never beat the funded project. Going from 8B to 14B barely helps.

2

More randomness is not more creativity. With the same RL training, raising the temperature reaches 19.8% originality. Ours reaches 56.3%.

3

RL with the same actions (ReAct) helps a lot, and creativity levels add about 10 more originality points. Removing Spark and Debate (Search + Write) costs 23 points.

4

Our 14B model is best on both, gaining 32.0 citation points and 66.2 originality points over Qwen3-14B. Scale helps once the model has learned when to be creative.

Head to head, Night-8B also beats GPT-4.1 on 75% of citation and 87% of originality comparisons.

(b) How varied are the ideas?

MethodParadigmContribution
1Qwen3-8B zero-shot0.7380.370
GPT-4.1 zero-shot0.7850.439
GIANTS-4B (RL)0.7730.347
Qwen3 + Temp (RL)0.7920.339
Qwen3 + ReAct (RL)0.8880.390
2Night-8B (Rpo)0.8530.362
  + no random swaps0.8440.358
1Night-8B (Rout)0.9430.425
How to read it. We label every proposal in two ways (defined below) and measure how evenly the proposals spread across seven categories. 1.0 means perfectly even; low means most proposals fall in one bucket.
1

Night-8B covers the widest range of research moves: 0.943, up from 0.738 for its base model, which is about 1.4 more categories.

2

The random swaps early in training help. Removing them from the process-reward model lowers diversity on both labels: 0.853 → 0.844 for paradigms and 0.362 → 0.358 for contributions.

What do the two labels mean?

A paradigm is the research move behind the idea: how it turns an opening into a direction (categories from Chen et al., 2026). A contribution type is what the work delivers in the end. They are independent: a new tool (artifact) could come from making a method more robust or from unifying two fields.

Paradigm: the research move

Synthesis
Bridge or unify separate fields, theories or methods.
Scope extension
Make prior work hold under weaker assumptions or new settings.
Robustification
Reduce failures, bias, risk or unreliability.
Formal derivation
A formal model, theorem, bound or taxonomy.
Empirical mapping
Systematic measurement, benchmarks or comparisons.
Artifact / system
Build a concrete system, tool or prototype.
Optimization / search
Tune, search or scale to find a better solution.

Contribution type: what it delivers

Empirical
New findings from data.
Artifact
A new system, tool, process or intervention.
Method
A reusable way of doing research or practice.
Theory
Reusable concepts, models, principles or frameworks.
Benchmark / data
A reusable dataset or evaluation resource.
Survey
A synthesis of existing work.
Opinion
An evidence-based position meant to shift the discussion.

Paradigms. The base model mostly proposes “build a system”: 44% artifact. Night-8B (O) cuts that to 14% and spreads out into synthesis (21%), scope extension (19%) and formal work (13%).

Stacked bar chart of primary research-idea paradigms for each method. Qwen-8B zero-shot is 44% artifact; Night-8B (O) is 21% synthesis, 19% scope extension, 9% robustification, 13% formal, 7% empirical mapping, 14% artifact and 16% optimization.

Contribution types. Every model leans toward artifacts, but Night-8B (O) leans least: 42% versus 68% for the base model, with more methods (25%) and theory (33%).

Stacked bar chart of primary contribution types for each method. Qwen-8B zero-shot is 68% artifact, 15% method, 13% theory; Night-8B (O) is 42% artifact, 25% method and 33% theory.

(O) is the outcome-only reward, Rout. (PO) adds the process reward, Rpo.

(c) Rewarding the final proposal vs. every step

MethodRewardCitation %Originality %
2ReActRout24.9446.60
Rpo2.106.54
2Temp.Rout11.5219.82
Rpo10.3513.33
1Night-8BRout29.8956.32
Rpo26.2061.61
How to read it. Rout scores only the final proposal. Rpo also scores each step along the way.
1

Scoring every step makes Night-8B more original (+5.3 points), but less likely to be cited (−3.7) and less varied.

2

For ReAct and temperature, the same reward hurts both scores. It only pays off when actions have meaningful creativity levels, so we use Rout by default.

Qualitative example

Same problem, four proposals

Every method got the same NSF topic: “Using AI to Transform Online Video Lectures into Effective and Inclusive Agent-Based Presentations.” The baselines polish how lectures are delivered. Our models rethink what a lecture is.

GPT-4.1Recombines existing tools
ReActAdds one new artifact
Night-8B (Rout)Reframes the interaction
Night-8B (Rpo)Most novel, least grounded
← Adapts existing lecture delivery Changes the learning interaction itself →
Hover or tap a highlighted phrase for an explanation. green novel or well-specified red incremental, vague or questionable

GPT-4.1

Zero-shot

Segment video lectures into topics, then have conversational agents re-present each segmentWhy red: incrementalLecture segmentation and chatbot tutors both already exist. Putting them together is integration work, not a new mechanism., with adaptive pacing, clarifications and multimodal deliveryWhy red: a feature listSign language, audio descriptions and adaptive pacing are valuable for accessibility, but they are well-known features. None of them sets this proposal apart from prior systems.. Evaluate against standard video lectures.

+ Well-structured four-phase plan that covers accessibility.
− No technical contribution beyond combining existing pieces.

Takeaway: safe and sensible, but close to what already exists.

ReAct

RL, same actions, no creativity levels

Adapt lectures in real time using a multimodal engagement scoreWhy green: a concrete new artifactOne score that fuses visual, auditory and physiological signals. It’s specific, measurable, and something later phases can build on. built from visual, auditory and physiological signals, fed into an RL engine that tunes pacing and content granularityWhy red: underspecifiedThe state, action and reward spaces are never defined, so it’s unclear how engagement signals actually become adaptation decisions.. Validated with a 120-person A/B study.

+ A concrete novel artifact and a controlled experiment.
− The core method is vague relative to the elaborate evaluation.

Takeaway: one new building block inside a fairly standard plan.

Night-8B (Rout)

Ours · outcome reward

Propose Agentify, which turns passive lectures into live, agent-mediated sessionsWhy green: a reframingRather than optimizing a recording after the fact, the lecture becomes a live, co-created session where an AI interleaves questions, hints and dialogue based on learner behavior.. A study randomizes 200 learners across 40 lectures into passive, over-moderated, fixed-tier and adaptive-tier groups, and includes a forced speed slider with eight settings (1×–8×)Why red: contrived detailThe rationale for this component is weak, and the related idea of co-evolving agent templates with “cognitive tiers” is underdeveloped..

+ Novel framing, testable hypothesis, sensible multi-group design.
− A few components are loosely motivated.

Takeaway: asks “what if a lecture were a conversation?” Bold, but still testable.

Night-8B (Rpo)

Ours · process + outcome

Propose Progressive Content Enactors: agents that deliberately embed controlled errors and ambiguityWhy green: the most novel directionPurposeful mistakes become a teaching tool, much as working through a flawed proof teaches more than reading a perfect one. It echoes error-based learning and productive failure, yet is clearly distinct from prior lecture systems. so learners must detect and resolve them, moving from passive reception to active problem solving. The background cites an engagement score of c = 0.12, +3.5σ retention and a 2.78σ efficacy gainWhy red: overreachNone of these numbers is independently verifiable or tied to a standard evaluation framework, which undermines the proposal’s credibility..

+ A fresh, actionable concept clearly differentiated from prior work.
− Unverifiable metrics and a weakly grounded evaluation.

Takeaway: the boldest idea and the least grounded. Scoring every step pushed originality at the cost of rigor.

Reading the comparison

Baselines improve the recording. GPT-4.1 and ReAct add better segmentation, pacing or sensing. The lecture is still something you watch.

Ours changes the premise. Both variants drop the idea that a lecture is passive. It becomes a conversation, or a puzzle to solve.

Bolder isn’t always better. The step-scored variant is the most original and the least rigorous. That is why outcome-only training is our default.

What the model learned

Higher levels really change behavior

When the model picks a higher creativity level, its output should get more varied and move further from the current draft. With our written-out levels, that link gets stronger during training. With temperature levels, it breaks down.

Line charts over training steps of the correlation between chosen creativity level and output entropy and semantic similarity, for AI Night-Scientist and the temperature variant.
Correlation between the chosen level and output behavior during training. Higher is better.

Scoring every step leads to more Spark

GPT-4.1 usually just searches, writes and stops. Night-8B trained with step rewards picks Spark (challenging an assumption) more often, usually at mid-to-high levels. Debate is rare, likely because it uses up more of the five-step budget.

Stacked bar charts of test-time action frequencies across five reasoning turns for GPT-4.1, Qwen3-8B-Base and AI Night-Scientist-8B.
Which actions each model picks at each of the five steps, at test time.

Conclusion

Key takeaways

01

Creativity can be learned

RL teaches the model when and how to leave the obvious path. The 14B model gains 32.0 citation points and 66.2 originality points over its base model.

02

Say what kind of creativity you want

Creativity levels written out in words beat simply raising the temperature. Randomness alone does not produce useful new ideas.

03

Broader ideas, not just bolder ones

Proposals spread across more kinds of research moves (+27.8%) and contributions (+14.9%), instead of mostly “build a system.”

04

Rewards shape how the model explores

Scoring every step makes the model reach for Spark more often and gives the most original proposals. Scoring only the final proposal gives the highest citation impact.

AI Night-Scientist is meant to be a creative collaborator for researchers, not a replacement. Its proposals are starting points that still need expert judgment and testing.

Citation

@article{kargupta2026nightscientist,
  title   = {Reinforcing Agentic Creativity in Scientific Ideation with Night Science},
  author  = {Kargupta, Priyanka and Cucerzan, Silviu and Mahajan, Shweti and Herring, Allen and Han, Jiawei and White, Ryen W. and Jauhar, Sujay Kumar},
  journal = {arXiv preprint arXiv:2609.35706},
  year    = {2026}
}