{
  "schema_version": "1.1",
  "id": "archive:https://arxiv.org/abs/2608.07460v1",
  "slug": "2608-07460v1-0tub7de",
  "url": "https://feed7.dev/p/2608-07460v1-0tub7de",
  "title": "CreativeInstruct: Scalably Teaching LLMs to Balance Quality, Creativity, and Diversity",
  "why_included": "CreativeInstruct adds learned control spans that recover base-model-like diversity after post-training, with reported gains in human creativity ratings and downstream RL training.",
  "summary": "CreativeInstruct teaches a model to insert **[StartCreativity] spans** that steer selected generations toward greater variation. The paper also proposes graph-edit distance for structural narrative diversity and reports a human preference for its creativity in **70.3% of cases** versus post-trained models.",
  "practical_implication": "For builders working on ideation, synthetic data, or exploratory agent behavior, this suggests creativity can be exposed as a learned generation mode instead of requiring several models at inference. The reported GRPO runs also improved by **about 4% on AMC** and **5 percentage points on MATH** over training from the post-trained checkpoint.",
  "agent_context": "CreativeInstruct teaches a model to insert **[StartCreativity] spans** that steer selected generations toward greater variation. The paper also proposes graph-edit distance for structural narrative diversity and reports a human preference for its creativity in **70.3% of cases** versus post-trained models.\n\nFor builders working on ideation, synthetic data, or exploratory agent behavior, this suggests creativity can be exposed as a learned generation mode instead of requiring several models at inference. The reported GRPO runs also improved by **about 4% on AMC** and **5 percentage points on MATH** over training from the post-trained checkpoint.\n\nThe evidence comes from narrative generation and two math benchmarks, not coding-agent tasks. The abstract does not establish how reliably the control spans transfer across domains, prompts, model families, or production constraints.",
  "source": {
    "name": "arXiv",
    "url": "https://arxiv.org/abs/2608.07460v1",
    "published_at": "2026-08-07T17:55:48.000Z"
  },
  "source_class": "blog_post",
  "content_type": "Paper",
  "layer": "model",
  "domains": [
    "research"
  ],
  "topics": [
    "reasoning"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The evidence comes from narrative generation and two math benchmarks, not coding-agent tasks. The abstract does not establish how reliably the control spans transfer across domains, prompts, model families, or production constraints."
  ],
  "connected_context": {
    "meaning": "This introduces creativity as an explicit learned generation mode rather than an inference-time ensemble or prompt-only tactic, with structural diversity measured separately from preference. The reported math gains suggest the training signal may affect reasoning as well as narrative variation, but the supplied evidence narrows adoption to experimentation: transfer to coding, other model families, and production constraints remains unestablished.",
    "corpus_size": 409,
    "generated_at": "2026-08-10T10:05:29.932Z",
    "connections": [
      {
        "title": "GradCuit: Credit-Assigned Gradient Flow Enables Robust and Interpretable Test-Time Latent Reasoning",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.02585v1",
        "feed7_url": "https://feed7.dev/p/2608-02585v1-1t870md",
        "reason": "CreativeInstruct learns a persistent controllable mode during training, whereas GradCuit adapts frozen-model latent states per query; together they distinguish weight-level behavior controls from test-time reasoning optimization."
      },
      {
        "title": "Data Quality Is the Compute Multiplier — Ari Morcos, DatologyAI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=_PdK6x7PQNM",
        "feed7_url": "https://feed7.dev/p/data-quality-is-the-compute-multiplier-ari-morcos-datologyai-0x7k2ve",
        "reason": "The data-quality argument supports treating CreativeInstruct’s curated creativity signal and task mixture as potential sources of the gains, while reinforcing that results from narrative and math data should not be assumed to transfer unchanged."
      },
      {
        "title": "$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2607.28582v1",
        "feed7_url": "https://feed7.dev/p/2607-28582v1-0egi1xh",
        "reason": "Both alter post-training behavior without requiring multiple inference models, but β-OPSD tunes teacher–reference regularization for reasoning stability while CreativeInstruct learns an explicit span-triggered creativity mode."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": "2026-08-07T17:55:48.000Z",
  "modified_at": "2026-08-07T17:55:48.000Z",
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/2608-07460v1-0tub7de",
    "json": "https://feed7.dev/p/2608-07460v1-0tub7de.json",
    "markdown": "https://feed7.dev/p/2608-07460v1-0tub7de.md"
  }
}