{
  "schema_version": "1.1",
  "id": "archive:https://github.com/MakazhanAlpamys/Soup",
  "slug": "soup-1rqg9hp",
  "url": "https://feed7.dev/p/soup-1rqg9hp",
  "title": "MakazhanAlpamys/Soup",
  "why_included": "Soup v0.73.2 repairs misleading release-gate scores, adds over-refusal checks and measures GPU noise floors; its layer streaming also puts 8B fine-tuning within 4 GB VRAM.",
  "summary": "Soup trains from YAML and can stream a frozen base one layer at a time. On an RTX 3050 Laptop, Llama 3.1 8B with NF4 used **3.32 GB peak VRAM**; the reported 119.6 tok/s predates a later correctness repair and has not been rerun there.",
  "practical_implication": "For release decisions, **v0.73.2** fixes two misranked suites, adds mini_over_refusal to catch models that reject benign requests, and introduces **--noise-floor N** to rerun the base and discount deltas within observed GPU variation. Builders should treat these checks as part of the training recipe, not rely on a single aggregate score.",
  "agent_context": "Soup trains from YAML and can stream a frozen base one layer at a time. On an RTX 3050 Laptop, Llama 3.1 8B with NF4 used **3.32 GB peak VRAM**; the reported 119.6 tok/s predates a later correctness repair and has not been rerun there.\n\nFor release decisions, **v0.73.2** fixes two misranked suites, adds mini_over_refusal to catch models that reject benign requests, and introduces **--noise-floor N** to rerun the base and discount deltas within observed GPU variation. Builders should treat these checks as part of the training recipe, not rely on a single aggregate score.\n\nThe project says the noise floor sizes observed variation but does not choose a significance threshold. Layer streaming also trades memory for repeated weight reads, and Python support is limited to **3.10–3.12**.",
  "source": {
    "name": "GitHub",
    "url": "https://github.com/MakazhanAlpamys/Soup",
    "published_at": null
  },
  "source_class": "tool",
  "content_type": "GitHub Repo",
  "layer": "benchmark",
  "domains": [
    "data"
  ],
  "topics": [
    "benchmark-integrity",
    "agent-evals",
    "agent-reliability"
  ],
  "verification": {
    "status": "needs_review",
    "label": "Needs Review",
    "method": "unverified",
    "verified_at": null
  },
  "uncertainty": [
    "The project says the noise floor sizes observed variation but does not choose a significance threshold. Layer streaming also trades memory for repeated weight reads, and Python support is limited to **3.10–3.12**."
  ],
  "connected_context": {
    "meaning": "This makes evaluation integrity part of the training loop: corrected suite rankings, a benign-request refusal check, and repeated-base noise measurement all weaken reliance on one aggregate delta. It confirms that small gains need scrutiny, but does not define statistical significance; the headline throughput is also no longer decision-grade because it predates a correctness repair and was not rerun.",
    "corpus_size": 462,
    "generated_at": "2026-08-16T10:03:27.784Z",
    "connections": [
      {
        "title": "When Will The Benchmaxxing Plague End? — Nick Heiner, Surge AI",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=-npY6XjM8CQ",
        "feed7_url": "https://feed7.dev/p/when-will-the-benchmaxxing-plague-end-nick-heiner-surge-ai-178gqcg",
        "reason": "Soup operationalizes the warning that leaderboard gains can mislead by correcting misranked suites, separating refusal behavior, and discounting changes within observed run variation."
      },
      {
        "title": "onepot-Bench 0: towards lab-aware in silico chemistry benchmarks",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.02595v1",
        "feed7_url": "https://feed7.dev/p/2608-02595v1-0l7cc8k",
        "reason": "Both separate capability from safety-related refusal behavior; Soup’s benign-request suite specifically catches over-refusal, complementing onepot-Bench’s distinct safety-refusal evaluation."
      },
      {
        "title": "The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping",
        "source_name": "arXiv",
        "source_url": "https://arxiv.org/abs/2608.06361v1",
        "feed7_url": "https://feed7.dev/p/2608-06361v1-1n3dr85",
        "reason": "The video study shows aggregate improvements can hide failed intermediate behavior; Soup reaches a parallel implementation conclusion by requiring suite-level checks rather than trusting one combined score."
      },
      {
        "title": "Verifiable Environments for AI in Biology — Kenny Workman, LatchBio",
        "source_name": "AI Engineer",
        "source_url": "https://www.youtube.com/watch?v=3ZMUiFaQ3qg",
        "feed7_url": "https://feed7.dev/p/verifiable-environments-for-ai-in-biology-kenny-workman-latchbio-1vs6y66",
        "reason": "Both show that evaluator correctness must itself be validated: Soup repaired suite ranking errors, while the biology work found brittle deterministic graders can reject valid paths."
      }
    ]
  },
  "lifecycle": "Current",
  "published_at": null,
  "modified_at": null,
  "supersedes": [],
  "expires_at": null,
  "formats": {
    "html": "https://feed7.dev/p/soup-1rqg9hp",
    "json": "https://feed7.dev/p/soup-1rqg9hp.json",
    "markdown": "https://feed7.dev/p/soup-1rqg9hp.md"
  }
}