{"accounting":{"admittedCells":492,"cellsByArm":{"A-native-skill":205,"B-flat-claude-md":205,"C-no-instructions":82},"excludedCells":0,"expectedCells":492,"failedHostOracles":[{"host":"w5","noOpReward":"0","oracleReward":"0","taskId":"pddl-airport-planning"},{"host":"w6","noOpReward":"0","oracleReward":"0","taskId":"pddl-tpp-planning"}],"unavailableCells":0},"execution":{"agentHarness":{"heldConstantAcrossArms":true,"location":"host","name":"Claude Code"},"armConstruction":{"owner":"Colophon","reason":"BenchFlow's standard modes could not express the Skill, flattened CLAUDE.md, and no-instructions arms while keeping non-instruction resources constant.","transform":"jinn.demo1.claude-md-flatten@1"},"grading":{"location":"pinned task container","verifier":"upstream task verifier"},"source":{"benchmark":"SkillsBench","commit":"b63b7b2850226b6aa4fb5929a8c1ac7bc4d9a6af","preservedPackageParts":["curated Skill bundles","task environments","oracles","verifiers"],"release":"v1.1","upstreamRuntime":{"name":"BenchFlow","usedForOfficialCells":false,"version":"0.6.3"}}},"limitations":["This is one model and one run. It does not establish a general result about Skills, CLAUDE.md, SkillsBench as a whole, or other models.","The population was a flat 41 statically admitted tasks. The paired estimate uses the pre-declared 14-task informative subset: every control replicate scored zero and at least one instructed arm had a positive mean.","Four tasks had a nonzero control result and 23 left both instructed arms at zero. No task was selected or dropped because of its outcome.","Fourteen informative tasks do not meet the official confirmatory floor of 21 units in 13 independence clusters.","The agent ran on the host while grading ran in the pinned task container. The agent-side environment was the host interpreter, not the task image.","The same operator designed, ran, graded, and sealed this comparison. Local evidence makes the process inspectable and reproducible; it does not prove honesty against the run owner.","Two on-host task oracles failed. Both tasks remain in the fail-closed 492-cell denominator.","Distinct droplets and cell keys do not establish distinct real-world parties.","This report does not certify or rank either instruction-loading path, and it is not a publication that Skills do not work."],"manipulationCheck":{"abMeanPpm":177642,"cCells":82,"cFullPass":7,"cMeanPpm":85366,"upliftPpm":92276},"population":{"flatTasks":41,"funnel":[{"stage":"Statically admitted; all run","tasks":41},{"stage":"Control arm not identically zero","tasks":4},{"stage":"Both instructed arms remained at zero","tasks":23},{"stage":"Pre-declared informative subset","tasks":14}],"officialFloor":{"independenceClusters":13,"met":false,"units":21}},"provenance":{"analysisManifestSha256":"822b2f7469dc2e58a3e72eee32688614d296ba20fc381d9a074e3935a68622b3","benchmarkSha256":"b88a7d07c3d28a4e61145535145f06d0fd1b6e7323b43c0ea86cf018d229b370","cohortSha256":"f279e3fe7de1308f8177bea06d316b0f6c5d546900f75033396fe55a6cf0d001","declarationSha256":"a31405a150a66753273e7b645e5b1391265564c9f0d33df814e4af93bdeb7a7e","internalRunId":"skillsbench-v1.1-demo1-final","matrixSha256":"1bd2dc6cd59df4784e78fca37907000b2f102ec69448a9a233acc9ca7c834470"},"question":{"arms":[{"id":"A-native-skill","label":"Native Skill","replicatesPerTask":5},{"id":"B-flat-claude-md","label":"Root CLAUDE.md","replicatesPerTask":5},{"id":"C-no-instructions","label":"No instructions","replicatesPerTask":2}],"comparison":"A-native-skill minus B-flat-claude-md","instructionBytes":"identical-between-a-and-b"},"result":{"confidenceInterval95Ppm":{"lower":-223444,"upper":129159},"estimatePpm":-47143,"informativeTasks":14,"interpretation":"The point estimate slightly favors root CLAUDE.md. The 95% interval includes zero and effects in either direction.","methodStatement":"This method estimates an effect; it does not gate one. No verdict, threshold, or selection was registered.","unit":"ppm-of-reward"},"schema":"colophon.report-presentation/1","sealedAt":"2026-08-18T00:00:00.000Z","selfRunDisclosure":"The same operator designed, ran, graded, and sealed this comparison and is using it to show Colophon. The artifact makes the method and evidence checkable; it does not prove honesty against the run owner.","slug":"skill-vs-root-claude-md-haiku-4-5","subject":{"benchmark":{"commit":"b63b7b2850226b6aa4fb5929a8c1ac7bc4d9a6af","name":"SkillsBench","release":"v1.1"},"model":"claude-haiku-4-5-20251001"},"summary":"The same instruction bytes were loaded as a native Skill or root CLAUDE.md, with a no-instructions arm. On this model and run, the estimate slightly favored CLAUDE.md, but the interval includes zero and effects in either direction.","title":"Do you need a Skill, or is CLAUDE.md enough?","verification":{"bundleFormat":"benchmark-product-public-bundle/5","checks":["manifest","evidence-closure","artifact-integrity","signature-validity","matrix-rederivation","report-verification","claim-consistency"],"command":"npx @colophon-claims/verify@0.1 ./bundle","readerAvailability":"available","reportEnvelopeSha256":"c66199af9e86dd9701f40c0a9d1fd8648dde4e50821f99335ce4052cfbb93583"}}