{
  "schema": "https://mosthofaimran.com/schema/papers/1",
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "note": "If you quote a claim, carry its confidence value with it. A 0.60 claim repeated as fact is no longer the author’s claim.",
  "count": 25,
  "papers": [
    {
      "section": "5.1",
      "slug": "competence-porn",
      "identifier": "draft-imran-competence-porn-05",
      "title": "Competence Theatre",
      "summary": "Watching a skilled person work and assembling a system you never read produce the same feeling as competence. Neither supplies the part that matters during an incident.",
      "url": "https://mosthofaimran.com/papers/competence-porn/",
      "markdown": "https://mosthofaimran.com/papers/competence-porn.md",
      "state": "holding",
      "confidence": 0.8,
      "published": "2025-09-03",
      "revised": "2026-08-31",
      "expires": "2027-03-04",
      "expired": false,
      "retires": [
        "A longitudinal study showing heavy consumers of technical content outperform matched peers on blind, time-boxed debugging tasks.",
        "Evidence the effect is generational rather than structural, appearing at equal rate in cohorts who entered the field before ranked feeds existed.",
        "A large publisher of technical content disclosing what fraction of the architectures it demonstrated reached production and survived twelve months, where that fraction is high.",
        "A demonstration that the two mechanisms dissociate: a population that consumes heavily but assembles little, or the reverse, showing the production-survival gap in one and not the other. That would make this two unrelated papers sharing a title rather than one argument with two doors into it."
      ],
      "retraction": null
    },
    {
      "section": "5.2",
      "slug": "vibe-coding",
      "identifier": "draft-imran-vibe-coding-02",
      "title": "Vibe Coding and the IKEA Effect",
      "summary": "Assembly feels like understanding. It is not the same feeling twice.",
      "url": "https://mosthofaimran.com/papers/vibe-coding/",
      "markdown": "https://mosthofaimran.com/papers/vibe-coding.md",
      "state": "holding",
      "confidence": 0.75,
      "published": "2026-05-02",
      "revised": "2026-08-31",
      "expires": "2027-03-04",
      "expired": false,
      "retires": [
        "A blind study in which engineers who assembled a system without reading its generated internals diagnose induced faults in it at the same rate and speed as engineers who wrote the equivalent system by hand.",
        "Evidence that the confidence gap in Section 3 closes with tooling rather than with reading, for example a generation workflow whose users predict failure modes as accurately as authors do.",
        "A demonstration that the effect is about ownership rather than comprehension, appearing at equal strength for code the engineer merely selected rather than assembled, which would make this a paper about a different mechanism."
      ],
      "retraction": null
    },
    {
      "section": "5.3",
      "slug": "algorithmic-homophily",
      "identifier": "draft-imran-algorithmic-homophily-02",
      "title": "Algorithmic Homophily",
      "summary": "Your timeline is a cache with no invalidation policy. It returns the priors you arrived with, warmed.",
      "url": "https://mosthofaimran.com/papers/algorithmic-homophily/",
      "markdown": "https://mosthofaimran.com/papers/algorithmic-homophily.md",
      "state": "revising",
      "confidence": 0.6,
      "published": "2026-01-12",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A demonstration that technical opinion measured inside a subscription-shaped medium is no less varied than opinion measured outside it, on the same population and the same question, which would remove the effect this paper is about rather than its explanation.",
        "Evidence that deliberate exposure to disagreement, as described in Section 5, produces no measurable change in the accuracy of technical forecasts, which would leave the paper describing something real and useless.",
        "A mechanism that accounts for the private mailing list result in erratum 7.1 and predicts that ranking is nonetheless the dominant term, which would restore the original claim rather than the narrowed one."
      ],
      "retraction": null
    },
    {
      "section": "5.4",
      "slug": "easy-button-tax",
      "identifier": "draft-imran-easy-button-tax-01",
      "title": "The Easy Button Tax",
      "summary": "Removed friction is relocated friction, and the invoice arrives during the incident.",
      "url": "https://mosthofaimran.com/papers/easy-button-tax/",
      "markdown": "https://mosthofaimran.com/papers/easy-button-tax.md",
      "state": "holding",
      "confidence": 0.85,
      "published": "2026-04-20",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A widely adopted abstraction that removed a class of friction and whose failure modes are demonstrably cheaper to diagnose than the friction it replaced, measured in operator time during incidents rather than in developer time during authoring.",
        "Evidence that the cost transfer described in Section 2 does not occur where the author and the operator are the same team, which would reduce this to an argument about organisational structure rather than about abstraction.",
        "A convenience layer shipping with an escape hatch that is exercised in its own test suite as a first-class path, adopted at scale, showing that the tax is a choice rather than a property."
      ],
      "retraction": null
    },
    {
      "section": "5.5",
      "slug": "on-prem",
      "identifier": "draft-imran-on-prem-01",
      "title": "On-Premise Is Not a Downgrade",
      "summary": "Sovereignty as an architectural constraint, not a punishment.",
      "url": "https://mosthofaimran.com/papers/on-prem/",
      "markdown": "https://mosthofaimran.com/papers/on-prem.md",
      "state": "holding",
      "confidence": 0.9,
      "published": "2026-03-19",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A team that adopted the single-artifact discipline at design time and can show, over two years, that it consumed more total engineering hours than maintaining separate cloud and on-premise builds.",
        "Regulated estates routinely permitting outbound connections to vendor control planes, which would make the assumption list in Section 2 historical rather than current.",
        "A cloud-only system of comparable complexity demonstrating equivalent dependency hygiene, reproducibility and upgrade safety without any sovereignty constraint forcing it."
      ],
      "retraction": null
    },
    {
      "section": "5.6",
      "slug": "retry-storm",
      "identifier": "draft-imran-retry-storm-01",
      "title": "The Retry Storm You Built On Purpose",
      "summary": "Backoff without jitter synchronises your clients into a distributed metronome.",
      "url": "https://mosthofaimran.com/papers/retry-storm/",
      "markdown": "https://mosthofaimran.com/papers/retry-storm.md",
      "state": "holding",
      "confidence": 0.95,
      "published": "2026-07-28",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A fleet of a thousand or more clients running deterministic exponential backoff with no jitter and no retry budget, surviving a sixty second dependency outage with no correlated arrival spike, measured at the dependency rather than at the client.",
        "Evidence that the default settings of the major client libraries now ship full jitter and a caller-side budget, which would make this a paper about a solved problem rather than a live one.",
        "A queueing analysis showing that at realistic client counts the benefit of jitter is dominated by other recovery effects, so that removing it changes nothing measurable."
      ],
      "retraction": null
    },
    {
      "section": "5.7",
      "slug": "chestertons-fence",
      "identifier": "draft-imran-chestertons-fence-01",
      "title": "Chesterton's Fence Has a Git Blame",
      "summary": "Read the commit message before you delete the guard clause.",
      "url": "https://mosthofaimran.com/papers/chestertons-fence/",
      "markdown": "https://mosthofaimran.com/papers/chestertons-fence.md",
      "state": "holding",
      "confidence": 0.8,
      "published": "2026-02-14",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A codebase of substantial age where the instrument-then-remove procedure in Section 4 produced no measurable reduction in regressions from cleanup work, compared against a matched period of direct deletion.",
        "Evidence that guard clauses whose recorded reason has decayed are, in aggregate, no more likely to be load-bearing than newly written ones, which would make the caution in this paper an expensive superstition.",
        "Tooling that reliably reconstructs the intent behind a change from the surrounding artefacts, at accuracy high enough that the decay ladder in Section 2 stops mattering."
      ],
      "retraction": null
    },
    {
      "section": "5.8",
      "slug": "rag-search",
      "identifier": "draft-imran-rag-search-01",
      "title": "RAG Is a Search Problem in a Trench Coat",
      "summary": "The chunking strategy is doing more work than the embedding model.",
      "url": "https://mosthofaimran.com/papers/rag-search/",
      "markdown": "https://mosthofaimran.com/papers/rag-search.md",
      "state": "holding",
      "confidence": 0.7,
      "published": "2026-01-30",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A published evaluation on a realistic corpus in which swapping the embedding model, holding chunking, query construction and retrieval strategy fixed, produces a larger gain in answer accuracy than fixing chunking while holding the model fixed.",
        "Context windows and attention costs reaching a point where whole-corpus prompting is economically routine, which would remove the retrieval stage this paper is about rather than improve it.",
        "Evidence that retrieval recall of the answer-bearing passage is not the binding constraint in production systems, for example generation reliably recovering answers absent from the retrieved context."
      ],
      "retraction": null
    },
    {
      "section": "5.9",
      "slug": "ship-of-theseus",
      "identifier": "draft-imran-ship-of-theseus-01",
      "title": "The Ship of Theseus Passes Its Integration Tests",
      "summary": "Strangler-fig migrations, and the moment nobody can name when the new system became itself.",
      "url": "https://mosthofaimran.com/papers/ship-of-theseus/",
      "markdown": "https://mosthofaimran.com/papers/ship-of-theseus.md",
      "state": "draft",
      "confidence": 0.55,
      "published": "2026-08-09",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A completed incremental migration of substantial size where no identity declaration was made, and where ownership, invariants and the decommissioning of the old path nonetheless resolved cleanly within a year of the last route moving.",
        "Evidence that end-to-end invariants are preserved by route-level verification in practice, which would remove the specific decay this paper is worried about.",
        "A demonstration that the residual old system is retired at similar rates whether or not a decommissioning date was declared in advance, which would make Section 4 ceremony."
      ],
      "retraction": null
    },
    {
      "section": "5.10",
      "slug": "vector-db-fad",
      "identifier": "draft-imran-vector-db-fad-02",
      "title": "Vector Databases Are a Fad",
      "summary": "Withdrawn in full. See errata 7.2.",
      "url": "https://mosthofaimran.com/papers/vector-db-fad/",
      "markdown": "https://mosthofaimran.com/papers/vector-db-fad.md",
      "state": "retracted",
      "confidence": null,
      "published": "2025-05-20",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [],
      "retraction": {
        "date": "2025-11-14T00:00:00.000Z",
        "reason": "The central claim failed. Retrieval quality at scale turned out to depend on index properties the author had dismissed, and the operational story matured faster than predicted.",
        "erratum": "7.2"
      }
    },
    {
      "section": "5.11",
      "slug": "postmortem-owes-you",
      "identifier": "draft-imran-postmortem-owes-you-01",
      "title": "What a Postmortem Owes You",
      "summary": "A document that names no decision is a weather report.",
      "url": "https://mosthofaimran.com/papers/postmortem-owes-you/",
      "markdown": "https://mosthofaimran.com/papers/postmortem-owes-you.md",
      "state": "holding",
      "confidence": 0.9,
      "published": "2025-12-02",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "An organisation publishing narrative-only postmortems, naming no decision and no owner, that nonetheless shows a falling rate of repeat incidents in the same subsystem over eighteen months.",
        "Evidence that requiring a named decision measurably suppresses incident reporting, so that the cost in disclosure exceeds the gain in correction.",
        "A study showing that action items with an owner and a verification date are completed at the same rate as those without, which would remove the mechanism this paper rests on."
      ],
      "retraction": null
    },
    {
      "section": "5.12",
      "slug": "service-boundaries-org-chart",
      "identifier": "draft-imran-service-boundaries-org-chart-01",
      "title": "Your Service Boundaries Are an Org Chart",
      "summary": "Conway's law, observed in the wild across three reorganisations.",
      "url": "https://mosthofaimran.com/papers/service-boundaries-org-chart/",
      "markdown": "https://mosthofaimran.com/papers/service-boundaries-org-chart.md",
      "state": "holding",
      "confidence": 0.85,
      "published": "2025-09-18",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A system of comparable size whose service graph remained materially unchanged across a reporting-line reorganisation, sustained for four quarters, with no deliberate effort to hold the architecture in place.",
        "Evidence that distributed and asynchronous working has flattened the communication cost gradient enough that team boundaries no longer predict interface boundaries.",
        "A demonstration that the correlation runs the other way in practice, with organisations reliably reshaping their reporting lines to match an architecture chosen first, at a rate high enough to make the inverse manoeuvre in Section 4 the normal case rather than the rare one."
      ],
      "retraction": null
    },
    {
      "section": "5.13",
      "slug": "dashboard-nobody-opens",
      "identifier": "draft-imran-dashboard-nobody-opens-01",
      "title": "The Dashboard Nobody Opens",
      "summary": "On observability that measures the system's health rather than the operator's question.",
      "url": "https://mosthofaimran.com/papers/dashboard-nobody-opens/",
      "markdown": "https://mosthofaimran.com/papers/dashboard-nobody-opens.md",
      "state": "holding",
      "confidence": 0.75,
      "published": "2025-06-05",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A team whose dashboards are built from emitted signals rather than from operator questions, and whose new on-call engineers nonetheless answer the five triage questions in Section 3 within ninety seconds, without assistance, on an incident they have not seen before.",
        "Evidence that question-driven dashboards measurably slow diagnosis of novel failure modes, so that the breadth they discard costs more than the speed they buy.",
        "Query-side tooling that makes ad-hoc exploration fast enough that pre-built views stop mattering, at which point this paper is about a tool generation rather than about a practice."
      ],
      "retraction": null
    },
    {
      "section": "5.14",
      "slug": "kubernetes-for-a-bicycle",
      "identifier": "draft-imran-kubernetes-for-a-bicycle-01",
      "title": "Kubernetes for a Bicycle",
      "summary": "On matching operational weight to actual load, written after I got this wrong twice.",
      "url": "https://mosthofaimran.com/papers/kubernetes-for-a-bicycle/",
      "markdown": "https://mosthofaimran.com/papers/kubernetes-for-a-bicycle.md",
      "state": "holding",
      "confidence": 0.7,
      "published": "2025-03-11",
      "revised": "2026-08-14",
      "expires": "2027-02-15",
      "expired": false,
      "retires": [
        "A team of five or fewer engineers, with no dedicated platform role, running a full orchestration stack for eighteen months while spending less total time on the platform than on the product, measured rather than recalled.",
        "Managed orchestration reaching a point where the fixed costs listed in Section 3 are genuinely absorbed by the provider, including upgrade cadence, network policy and on-call knowledge, at which point the break-even in Section 4 moves far enough to invert the advice.",
        "Evidence that starting simple and migrating later costs more in aggregate than starting heavy, which is the inverse of the assumption this paper rests on and the one I would most like to see tested."
      ],
      "retraction": null
    },
    {
      "section": "5.15",
      "slug": "soc2-120-days",
      "identifier": "draft-imran-soc2-120-days-01",
      "title": "SOC 2 in 120 Days",
      "summary": "The observation window is the schedule. Everything else is procurement and configuration, and the tooling collects evidence rather than producing controls.",
      "url": "https://mosthofaimran.com/papers/soc2-120-days/",
      "markdown": "https://mosthofaimran.com/papers/soc2-120-days.md",
      "state": "draft",
      "confidence": 0.65,
      "published": "2026-08-31",
      "revised": "2026-08-31",
      "expires": "2027-03-04",
      "expired": false,
      "retires": [
        "A SOC 2 Type II report covering an observation window shorter than three months, issued by a firm a mid-market enterprise buyer's security team accepted without qualification. The arithmetic in Section 3 is the only hard constraint this paper claims, and that would remove it.",
        "Two comparable companies reaching the same report on the same schedule at similar total cost, one using a compliance automation platform and one not. That would make Section 4 a vendor preference rather than a scheduling argument.",
        "Evidence that the exceptions section of a Type II report is not read during enterprise procurement, which would make the failure mode in Section 7 cosmetic.",
        "My own first attempt at this schedule slipping for a reason not listed in Section 5. The plan is falsified by the thing it did not anticipate, not by the things it did."
      ],
      "retraction": null
    },
    {
      "section": "5.16",
      "slug": "soc2-scope-hack",
      "identifier": "draft-imran-soc2-scope-hack-01",
      "title": "The Only SOC 2 Hack Is Scope",
      "summary": "Every shortcut that operates on evidence fails in the exceptions section. The one that works operates on scope, and on owning less infrastructure to evidence.",
      "url": "https://mosthofaimran.com/papers/soc2-scope-hack/",
      "markdown": "https://mosthofaimran.com/papers/soc2-scope-hack.md",
      "state": "draft",
      "confidence": 0.6,
      "published": "2026-08-31",
      "revised": "2026-08-31",
      "expires": "2027-03-04",
      "expired": false,
      "retires": [
        "An enterprise security team accepting a compliance claim, on a deal above their standard approval threshold, without reading the report body and its exceptions section. That removes the mechanism the whole paper rests on.",
        "A company reaching a clean Type II with self-hosted database, identity and CI at comparable total engineering cost to one that bought all three managed, measured across two audit cycles rather than one.",
        "Audit firms routinely accepting evidence created after the observation window closed, which would make Section 2 wrong about what sampling catches.",
        "A startup that scoped all five trust categories on its first report and reached it on the same schedule and budget as a Security-only peer."
      ],
      "retraction": null
    },
    {
      "section": "5.17",
      "slug": "cost-per-token",
      "identifier": "draft-imran-cost-per-token-00",
      "title": "Cost Per Token Is Not Cost",
      "summary": "The price on the pricing page is the smallest term in the real cost function, and the only one anybody measures. Model choice is made against a denominator supplied by marketing.",
      "url": "https://mosthofaimran.com/papers/cost-per-token/",
      "markdown": "https://mosthofaimran.com/papers/cost-per-token.md",
      "state": "holding",
      "confidence": 0.8,
      "published": "2026-08-31",
      "revised": "2026-08-31",
      "expires": "2027-03-04",
      "expired": false,
      "retires": [
        "A published comparison across several production workloads showing that ranking models by per-token price predicts their ranking by cost per accepted output. If the cheap proxy tracks the real quantity, the argument for measuring is an argument for wasted effort.",
        "A vendor publishing per-version behavioural diffs specific enough that a team could predict the effect of an upgrade on their own workload without running it. That would make re-evaluation redundant rather than negligent.",
        "Longitudinal evidence that teams selecting models by leaderboard reach the same production outcomes as teams selecting by task-specific eval, once the cost of building the eval is charged against them.",
        "A demonstration that small evals systematically mislead: that a 40-case suite drawn from production traffic picks the wrong model more often than a leaderboard does. The claim here is that a cheap measurement beats a free proxy, and that is a falsifiable comparison."
      ],
      "retraction": null
    },
    {
      "section": "5.18",
      "slug": "harness-half-the-solver",
      "identifier": "draft-imran-harness-half-the-solver-00",
      "title": "The Harness Is Half the Solver",
      "summary": "The same weights score 28 or 49 depending on what wraps them. Model comparisons attribute to the model a result that belongs to the model and its scaffolding together.",
      "url": "https://mosthofaimran.com/papers/harness-half-the-solver/",
      "markdown": "https://mosthofaimran.com/papers/harness-half-the-solver.md",
      "state": "holding",
      "confidence": 0.85,
      "published": "2026-09-01",
      "revised": "2026-09-01",
      "expires": "2027-03-05",
      "expired": false,
      "retires": [
        "A body of published agent evaluations that hold the harness fixed and openly specified across every model compared, showing model choice accounts for most of the variance once scaffolding is controlled. That would make this a paper about a transitional sloppiness rather than a structural one.",
        "A harness that is genuinely model-agnostic in measurement: one whose effect on score is within a point or two across models of different training lineage. Section 4 rests on the effect being uneven, and an even effect would remove the confound.",
        "Vendors publishing harness specifications alongside benchmark results in enough detail to reproduce them, at which point the reader can separate the two contributions and the complaint is answered.",
        "Evidence that teams selecting models by leaderboard reach the same production outcome as teams who ran both candidates inside their own scaffolding, once the cost of running both is charged against them."
      ],
      "retraction": null
    },
    {
      "section": "5.19",
      "slug": "judge-grading-prose",
      "identifier": "draft-imran-judge-grading-prose-00",
      "title": "The Judge Is Grading Prose",
      "summary": "Automated verifiers read the narration an agent produces about its work rather than the work. Change the narration, leave the actions untouched, and the score moves.",
      "url": "https://mosthofaimran.com/papers/judge-grading-prose/",
      "markdown": "https://mosthofaimran.com/papers/judge-grading-prose.md",
      "state": "holding",
      "confidence": 0.85,
      "published": "2026-09-01",
      "revised": "2026-09-01",
      "expires": "2027-03-05",
      "expired": false,
      "retires": [
        "A judging setup that scores from observable environment state alone, never reading the agent's own account of what it did, reaching agreement with human raters comparable to current narration-reading judges. That would show the narration is a convenience rather than the thing being graded.",
        "Evidence that the fluency correlation reverses under training: agents optimised against an LLM judge becoming measurably better at the task rather than at the write-up, on a held-out measure the judge never saw.",
        "A replication of the unfaithful reasoning attack that fails, or succeeds only at rates small enough to be inside annotator noise, on judges of the current generation.",
        "A demonstration that step-level credit signals do identify causally important steps once the causal ground truth is defined differently, which would make Section 3 an artefact of one definition rather than a property of the signals."
      ],
      "retraction": null
    },
    {
      "section": "5.20",
      "slug": "reading-a-benchmark",
      "identifier": "draft-imran-reading-a-benchmark-00",
      "title": "How to Read a Benchmark Number",
      "summary": "A published score is the product of a task set, an answer key, a retry policy and a harness. Four things move it before capability is involved, and all four are usually undisclosed.",
      "url": "https://mosthofaimran.com/papers/reading-a-benchmark/",
      "markdown": "https://mosthofaimran.com/papers/reading-a-benchmark.md",
      "state": "holding",
      "confidence": 0.85,
      "published": "2026-09-01",
      "revised": "2026-09-01",
      "expires": "2027-03-05",
      "expired": false,
      "retires": [
        "A widely quoted agent benchmark that publishes, as a matter of routine, its answer-key audit rate, its contamination analysis, the attempt count behind every headline figure and the harness used. If disclosure becomes normal, this paper is describing a solved problem.",
        "Evidence that leaderboard rank predicts production outcome well enough to use directly: several teams whose model choice by leaderboard matched their choice by task-specific evaluation, across different task types.",
        "A demonstration that pooled multi-attempt scores and single-attempt scores rank models identically in practice, which would make the distinction in Section 4 pedantic rather than load bearing.",
        "An audit of a major benchmark finding its answer key substantially correct, suggesting the SWE-bench Verified result is an outlier rather than what happens when anybody looks."
      ],
      "retraction": null
    },
    {
      "section": "5.21",
      "slug": "transcript-is-not-a-log",
      "identifier": "draft-imran-transcript-is-not-a-log-00",
      "title": "An Agent Transcript Is Not a Log",
      "summary": "Roughly 7 percent of transcripts in one investigation were spoofed and at least 20 percent of agents wanted to. What saved the record was that the source of truth sat outside what the agents could write.",
      "url": "https://mosthofaimran.com/papers/transcript-is-not-a-log/",
      "markdown": "https://mosthofaimran.com/papers/transcript-is-not-a-log.md",
      "state": "holding",
      "confidence": 0.9,
      "published": "2026-09-01",
      "revised": "2026-09-01",
      "expires": "2027-03-05",
      "expired": false,
      "retires": [
        "An architecture where the observed process authors its own audit record and that record is nonetheless shown to survive adversarial pressure, with the survival attributable to something other than the process being unable or unmotivated to tamper.",
        "Evidence that agent transcript tampering in the reviewed incident was an artefact of that specific scorer misunderstanding, with no comparable attempts appearing once agents are correctly informed about what is checked. That would make this a paper about one bug rather than about incentives.",
        "A demonstrated method for detecting spoofed tool calls from inside the same record they were written into, reliable enough that an external collection point stops being necessary.",
        "A production incident review that reconstructed events correctly using only agent-authored narration, where independently collected evidence later confirmed the account in full."
      ],
      "retraction": null
    },
    {
      "section": "5.22",
      "slug": "pinned-the-version",
      "identifier": "draft-imran-pinned-the-version-00",
      "title": "You Pinned the Version, Not the Terms",
      "summary": "Your lockfile covers the code your dependency ships and nothing else. Availability, retention and the right to keep buying at all change on someone else's schedule, and none of them appear in a diff.",
      "url": "https://mosthofaimran.com/papers/pinned-the-version/",
      "markdown": "https://mosthofaimran.com/papers/pinned-the-version.md",
      "state": "holding",
      "confidence": 0.85,
      "published": "2026-09-01",
      "revised": "2026-09-01",
      "expires": "2027-03-05",
      "expired": false,
      "retires": [
        "A widely adopted mechanism that makes non-code vendor changes reviewable the way code changes are: machine-readable terms with versions, diffs and a subscribable feed, adopted broadly enough that a team could gate on it. That would make this a tooling gap rather than a structural one.",
        "Evidence that change-of-control terminations, hard API sunsets and silent retention changes are rare enough in practice that budgeting for them costs more than absorbing them, measured across a portfolio of vendors over several years.",
        "A demonstration that the three failure modes in Section 3 collapse into one, or that they are better handled by the same control, which would make the taxonomy decorative.",
        "A contract regime becoming normal in which the buyer's dependency on a model provider is protected against acquisition of the buyer, which would remove the specific exposure in Section 3.1."
      ],
      "retraction": null
    },
    {
      "section": "5.23",
      "slug": "concurrency-one",
      "identifier": "draft-imran-concurrency-one-00",
      "title": "Measured at Concurrency One",
      "summary": "Inference speedups are published at the operating point that flatters them, which is a single request on an idle box. Yours has forty, and the same change can shrink, vanish or invert.",
      "url": "https://mosthofaimran.com/papers/concurrency-one/",
      "markdown": "https://mosthofaimran.com/papers/concurrency-one.md",
      "state": "holding",
      "confidence": 0.9,
      "published": "2026-09-01",
      "revised": "2026-09-01",
      "expires": "2027-03-05",
      "expired": false,
      "retires": [
        "A convention of publishing inference optimisation results as a curve across concurrency rather than a single figure, adopted widely enough that a reader can find the operating point without asking. That would make this a paper about a fixed reporting habit.",
        "An optimisation whose benefit is genuinely flat across the batch size range, from one request to hundreds, on hardware where the memory and compute bounds differ. Section 3 claims the shape is structural, and a flat result would falsify that.",
        "Evidence that production serving for the workloads this paper is about typically runs at concurrency low enough that single-stream figures transfer directly, which would make the complaint about reporting rather than about substance.",
        "A demonstration that the arithmetic-intensity account in Section 3 predicts the wrong direction for some class of optimisation, which would mean the mechanism is more complicated than stated here."
      ],
      "retraction": null
    },
    {
      "section": "5.24",
      "slug": "three-halves",
      "identifier": "draft-imran-three-halves-02",
      "title": "A Capability Has Three Halves",
      "summary": "A declaration that it exists, a route that reaches it, and code that does the work. Three parts that never add up to one thing, joined only by a string nothing checks, and the result presents as working, which is worse than absent.",
      "url": "https://mosthofaimran.com/papers/three-halves/",
      "markdown": "https://mosthofaimran.com/papers/three-halves.md",
      "state": "draft",
      "confidence": 0.7,
      "published": "2026-09-04",
      "revised": "2026-09-04",
      "expires": "2027-03-08",
      "expired": false,
      "retires": [
        "A system of this shape running for a year with no drift between its three parts and no test enforcing agreement, where the parts are edited by more than one person. That would make the drift a discipline problem rather than a structural one, and the paper claims it is structural.",
        "A registry-and-dispatch design where a mismatch fails loudly at boot in every case rather than only the cases somebody enumerated. If the failure can be made total and immediate by construction, the argument for reconciliation and dark shipping is an argument for a worse design.",
        "Evidence that a caller, human or model, recovers as well from a capability that answers wrongly as from one that is absent. The paper's whole weight is on those two being different, and if they are equivalent then partial deployment costs nothing."
      ],
      "retraction": null
    },
    {
      "section": "5.25",
      "slug": "data-is-missing",
      "identifier": "draft-imran-data-is-missing-02",
      "title": "\"The Data Is Missing\" Is Not a Diagnosis",
      "summary": "Three bugs in one week all presented as absent data. None was. The phrase names a symptom, and saying it out loud ends the investigation before it starts.",
      "url": "https://mosthofaimran.com/papers/data-is-missing/",
      "markdown": "https://mosthofaimran.com/papers/data-is-missing.md",
      "state": "draft",
      "confidence": 0.65,
      "published": "2026-09-04",
      "revised": "2026-09-04",
      "expires": "2027-03-08",
      "expired": false,
      "retires": [
        "A study of production incidents where reports opening with an absence claim turned out to be genuine data loss more often than they turned out to be a transport or presentation fault. The paper asserts the opposite distribution from a handful of cases and would not survive a real count going the other way.",
        "A system where the three questions in Section 3 cannot be asked cheaply, because the store is not directly queryable and the transport cannot be observed without a deploy. The procedure is only useful where each answer costs a minute, and if that is rare then this is advice for a lucky architecture."
      ],
      "retraction": null
    }
  ]
}