← Stats Desk / API
Tokens

Drive Stats Desk from your own code

A reporting review, not a reanalysis. The model reads your text; it never sees raw data, never runs a test and never computes a p value itself. Your manuscript travels in text (up to 24,000 characters of numbered lines; longer text keeps the head and the tail and drops the middle with a marker), so leave out anything you may not share, such as unpublished participant details, before you send it. It is not medical, regulatory or clinical-trial advice.

Everything the web page does is available over HTTP. Send a Statistical analysis subsection, Results paragraphs, figure legends or reviewer comments and get the same review back: a verdict (not_ready, needs_revision, ready), a one-sentence headline, a short TL;DR, issues ranked P0 to P2 that each cite a line, quote it and say what a statistical reviewer would challenge and how to fix it, the lane's body (the scope, a claim map, a ready-to-paste revision and the reviewer risk for an audit; the drafted section, reporting notes and a legend template for a draft), the factual questions only the authors can answer, and one answer per browser flag. The natural use is a pre-submission check in a writing pipeline: run the statistics text through here and hold the manuscript on not_ready.

The model does not do the arithmetic. Every reported t, F, chi-square, r and z statistic with its degrees of freedom is found first by the page's free prescan in statkit.js, which turns it back into the p value it implies (allowing for the rounding of both numbers, with a one-tailed allowance and a decision-error test at 0.05) and marks it consistent, inconsistent, decision_error and so on. The same prescan flags n counted as cells or spines, missing multiple-comparison corrections, undefined error bars, threshold-only p values, trend language and more. The model is told to quote those computed values by id (T1) and never to state a p value of its own. The result is sent as facts, a JSON string. See the facts string.

Two lanes: the task field

Every request names its lane in task. Stats Desk has two, with two different reply bodies.

taskwhat it is forlane body in the reply
auditReview the statistical reporting of the supplied text: the independent unit and how n relates to it, paired or nested structure, multiplicity, reported p values against their statistics, effect sizes and uncertainty, error-bar definitions, exclusions, software, tails, overclaiming. Returns issues ranked P0/P1/P2, a map of the central claims, and ready-to-paste revised text for the part that most needs it, with AUTHOR_INPUT_NEEDED wherever a fact is missing.scope, claims, revision, reviewer_risk.
draftDraft a ready-to-paste Statistical analysis subsection from the text, the authors' design notes (design) and, if given, an earlier audit (audit), in a fixed order: software; summary convention; sample size and replication; tests or models per comparison; multiplicity; exact reporting and thresholds; exclusions. A method the audit recommends but the authors have not confirmed is written as AUTHOR_INPUT_NEEDED, never as done. The verdict describes the DRAFTED section, and issues lists only what the draft cannot resolve on its own.draft_text, reporting_notes, legend_template.

Both lanes share one envelope (lane, verdict, headline, tldr, issues, author_input_needed, prescan_responses); see the output contract. If task is missing or unrecognised, the model chooses the closest lane, answers with that lane's contract (never a blend) and names it in lane. Do not rely on that: always send "audit" or "draft". The page's own guard (StatKit.mustBeObject) refuses any other value, and its checks flag a reply whose lane differs from the task asked.

The prescan is the same for both lanes: facts depends only on the text, the target and the design notes. The same text run in two lanes is still two different runs, with two different Idempotency-Keys. The page's own flow is audit first, then draft: its "Draft the section from this audit" button sends the audit's issues and author questions as plain text (Recon.auditText, one line per issue and one per question) in audit, with the same text and design notes.

Input fields

The body is one flat JSON object, built in the page by StatKit.buildInput. Every value is a string. Required: task, text, facts.

fieldtyperequiredwhat it holds
taskstringyes"audit" or "draft"; see the lanes.
targetstringno"general", "nature" (the flagship journal Nature) or "other". Anything else is treated as general by the prescan. Only with nature does the model say "Nature asks for", and only for Nature's initial-submission statistical items (tests and tails named, error bars defined, exact n, repeat counts for representative results, exact P values, F and t with their degrees of freedom, replicates defined); for other targets it calls these good reporting practice. The prescan also raises the severity of tails_unstated, df_missing, test_stat_missing and n_range to medium for Nature, and adds repeats_unstated.
journalstringnoThe journal's name when target is other, else "". Cut at 120 characters.
titlestringnoA label for the manuscript ("Feedback timing and workload"). Context only. Cut at 200 characters.
designstringnoThe authors' own notes on the design: groups, units, replicates, software, what was or was not corrected, why samples were excluded. These count as facts just like the text, so a number or method stated here may appear in the drafted text. Cut at 3,000 characters on a word, the cut marked [... the rest of the design notes (N characters) was cut ...]. Most useful in the draft lane.
textstringyesThe manuscript text with every line prefixed by its number and "| " ("12| Spine density was higher ..."), so each issue can cite a real line. At most 24,000 characters: longer text keeps about 60% of the budget from the head and the rest from the tail, on whole lines, with a marker line in between: [... lines A-B were not sent (N lines); the facts were computed from the whole text ...]. Trailing blank lines are dropped before numbering.
factsstringyesA JSON string (the output of JSON.stringify), never an object. In the page it is the browser's free prescan of the WHOLE text, including any clipped middle; its keys are listed below.
auditstringdraft onlyThe issues from an earlier audit, as plain text, one per line ("I1 [S1] P0 stat_decision_error at line 10: ... Fix: ...", then "Author input needed: ..." lines). Cut at 8,000 characters. May be "": the model then drafts from the text, the design notes and the flags. The page sends this key only in the draft lane.
questionstringnoYour own question, answered inside tldr as a bullet starting "Answer:". Cut at 1,200 characters on a word; the page sends "" when empty.
retry_notestringnoOnly on a reformat retry, after a reply that could not be parsed: say what was wrong. The page adds it to the body of the second attempt and never on a first run. See truncation.

The facts string

In the web page, facts is computed by the browser before you pay for anything: the free prescan (statkit.js) reads the whole text, recomputes every test result it can parse, raises the flags you see on the page, lists what the text names, and serializes the result. An API caller has two options: build the same object with statkit.js, which runs unchanged in Node (see building the body), or send a minimal one yourself.

keywhat it holds
targetThe target as sent, with the journal in brackets for other ("other (eLife)").
browser_verdictThe prescan's hint: not_ready (any high flag), needs_revision (any medium), else ready. The model's verdict is never looser unless it dismissed the flags that set it.
flagsUp to 24 flags, worst first: id (S1, S2, ...), severity (high, medium, low), priority (P0, P1, P2, one to one with severity), category, message and lines (up to 12 line numbers; [] for an absence). The categories are in the next table.
statsEvery test result the browser found, as {id, line, test, df1, df2, value, reported, p_reported, p_computed, p_range, status}. id is T1, T2, ...; test is t, F, chi2, r or z; df1/df2 are null when absent; reported is the matched text (up to 140 characters); p_reported is "p = .034", "p < 0.01", "ns" or "none". p_computed is the p value the statistic implies (two-tailed for t, z and r) to four significant figures, and p_range is [low, high] allowing for the rounding of the printed statistic; both are null when status is impossible. These are the ONLY computed p values the model may quote.
n_mentionsUp to 40 sample sizes found as n = ...: {line, text, unit, unit_word, range}. unit is independent (mice, patients, participants, cultures, independent experiments...), subsample (cells, spines, fields, wells, images, technical replicates...) or unclear; unit_word is the word that followed; range is true for n = 3-4.
tests_named, corrections_named, software_named, summary_conventionsWhat the text names: tests ("t-test (paired)", "one-way ANOVA", "mixed-effects model", ...), multiple-comparison corrections ("Bonferroni", "Holm", "Benjamini-Hochberg / FDR", "Tukey", ...), software ("R", "GraphPad Prism", "SPSS", ...) and summary conventions ("mean ± s.d.", "mean ± s.e.m.", "confidence interval", "median and IQR", "box plot convention").
named_in_design_notes{tests, corrections, software}: the same lists read from design.
tails_stated, one_tailed_declaredWhether the text says one- or two-tailed (or sided) anywhere, and whether it says one-tailed. With a declared one-tailed test, a statistic that matches its p only as a one-tailed value counts as consistent.
p_counts{exact, threshold, ns}: how many p values are printed as p = x, as p < x, and as ns / not significant.
clippednull, or {total_lines, dropped} (dropped as "A-B") when the middle of the text was cut from text.
stats[].statusmeaning
consistentThe reported p lies within what the statistic implies, allowing for rounding of both numbers (or as a one-tailed p when the text declares one-tailed tests).
inconsistentIt does not, but both are on the same side of 0.05 (flag stat_inconsistent, medium).
decision_errorIt does not, and the reported p is on the other side of 0.05 from the implied one: a significant result that is not, or the reverse (flag stat_decision_error, high).
one_tailed_onlyA t, z or r result that matches only if the p is one-tailed, and the text never says one-tailed tests were used (flag stat_one_tailed_only, medium).
impossibleThe value cannot occur: r beyond ±1, a negative F or chi-square, a missing df where one is required, or p above 1 (flag stat_impossible, high).
no_pA statistic with no p value next to it; p_computed is still given.
flag categoryseveritywhat the prescan matched
stat_impossiblehighA stats entry with status impossible.
stat_decision_errorhighA stats entry with status decision_error.
stat_inconsistentmediumA stats entry with status inconsistent.
stat_one_tailed_onlymediumA stats entry with status one_tailed_only.
n_subsample_unithigh, or low when the text names a mixed, nested or per-animal averaging approachn counts cells, spines, fields, images or technical replicates (possible pseudoreplication).
multiple_comparisonshighSix or more p values or test results, or wording about many comparisons, with no correction or declared family named.
p_zeromediumA p value printed as zero (p = 0.000).
p_threshold_onlymediumEvery p value is a threshold; no exact values.
trend_languagemedium"Marginally significant", "a trend toward significance".
significance_differencemediumOne effect significant, another not, read as the effects differing.
summary_undefinedmediumError bars or ± values never defined as s.d., s.e.m. or a CI.
exclusions_vaguemediumExclusions or outliers without the rule, its timing or the count.
stars_undefinedmediumSignificance stars whose thresholds are never defined.
n_missingmediumTest results but no sample size anywhere.
repeats_unstatedmedium (Nature only)Representative results with no repeat count.
tails_unstatedmedium for Nature, else lowt, z, r or rank tests with no statement of tails.
df_missingmedium for Nature, else lowA t or F printed without degrees of freedom.
test_stat_missingmedium for Nature, else lowp values for t-tests or ANOVAs with no t or F statistic at all.
n_rangemedium for Nature, else lown given as a range (n = 3-4).
ns_without_plow"ns" or "not significant" with no p value or estimate.
overclaimlow"Highly significant", "proves", "conclusively".
association_as_causelowAn association written in causal or mechanistic terms.
randomization_blindinglowAnimals or participants compared with no word on randomization or blinding.
sem_usedlows.e.m. used as the summary convention.
normality_small_nlowNormality asserted or tested with n of 6 or fewer per group.
software_missinglowInference reported with no statistical software named.

The flags drive the reply. Every flag must come back exactly once in prescan_responses (confirmed or dismissed, with the reason and the line), and every confirmed high or medium flag must be carried by an issue with that flag's id in ref. They are pattern matches and can be wrong: the model is told to dismiss honestly, for example n counting cells when a mixed model with animal as a random effect is described, or a one_tailed_only stat where the Methods do state one-tailed tests.

Sending facts without the prescan. Empty lists are allowed. A minimal facts, as the JSON string you put in the body:

"facts": "{\"flags\":[],\"stats\":[]}"

It costs you the reconciliation the page does. With no flags there is nothing to confirm or dismiss, prescan_responses comes back [], the verdict has no floor, and the issues rest on the model's own reading of the text. With no stats, the model has no browser-computed p value to cite, and it is told never to compute one, so it can say a p value looks doubtful but cannot tell you what the statistic implies. Your own checks lose their reference too. Running statkit.js is free and gets you all of it.

Building the body

The surest way to match the page is to run the page's own engine. Save statkit.js from this site (it runs unchanged in Node via require()) and give analyze the same fields the page's form has; buildInput then numbers and clips the text, cuts the text fields and serializes the facts:

set fieldwhat it holds
laneaudit or draft (anything else is analyzed as audit). Becomes task, and decides whether audit is sent.
target, journalAs in the input fields.
titleThe manuscript label.
designThe design notes, raw.
textThe raw text, without line numbers.
// make-body.js - build the run body with the SAME engine the web page uses.
// Save statkit.js from https://stats-desk.skillsafe.ai/ next to this file.
// Usage: node make-body.js survey.txt audit general
//        node make-body.js mouse.txt draft nature design.txt audit.txt
const fs = require("fs");
const K = require("./statkit.js");

const [file = "survey.txt", lane = "audit", target = "general", designFile, auditFile] = process.argv.slice(2);
const A = K.analyze({
  lane: lane,                          // audit or draft
  target: target,                      // general, nature or other
  journal: "",                         // the journal's name when target is other
  title: "Feedback timing and workload",
  design: designFile ? fs.readFileSync(designFile, "utf8") : "",
  text: fs.readFileSync(file, "utf8")  // raw text; buildInput numbers the lines and clips
});
const audit = lane === "draft" && auditFile ? fs.readFileSync(auditFile, "utf8") : "";
const body = K.buildInput(A, { question: "", audit: audit });
fs.writeFileSync("body.json", JSON.stringify(K.mustBeObject(body)));
console.log(A.flags.length, "flags,", A.stats.length, "stats; browser verdict", A.hint);
console.log("Idempotency-Key: stats-desk:" + body.task + ":" + K.hashInput(body) + ":a1");

Run on the texts from the page's examples, this produces the worked requests below, byte for byte: the psychology experiment (audit, target general, hash lb7cqsc4tqiq) and the neuroscience draft (draft, target nature, with design notes and an audit, hash 1cehxnmtrgdzo). Trailing blank lines are dropped before numbering, so a trailing newline does not change the hash; any other edit does. mustBeObject is the page's own guard: it throws unless the body is a plain object with task audit or draft, non-empty text and a facts string. From another language, send the same field names, number the lines yourself ("1| ", "2| ", ...) and build facts with the keys above, or the minimal one.

For the draft lane, turn an audit reply into the audit text the same way the page's "Draft the section from this audit" button does:

// audit-text.js - turn an audit reply into the draft lane's audit text, as the page does.
// Save recon.js next to statkit.js. Usage: node audit-text.js reply.json
const fs = require("fs");
const R = require("./recon.js");

const reply = R.normalize(R.parseResult(fs.readFileSync(process.argv[2] || "reply.json", "utf8")), "audit");
fs.writeFileSync("audit.txt", R.auditText(reply));
console.log(reply.issues.length, "issues and", reply.author_input_needed.length, "questions written to audit.txt");

Base URL and the envelope

Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses the same envelope, so one helper covers the whole API:

{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}

The token is minted for this app (the guest endpoint takes {"slug":"stats-desk"} in its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….

The input object IS the request body. There is no {"input": …} wrapper. A wrapped body is answered with an unknown field 'input' warning, and the model never sees your text.

Error codes

statuscodewhat to do
400validation_errorA field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object.
401unauthorizedThe token is missing, malformed or expired. Get a new one from the token page.
402payment_requiredThe balance is below min_credits. Call /estimate first and top up.
403forbiddenThe token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token.
404not_foundUnknown job id, or the app slug does not exist.
409conflictThe same Idempotency-Key was replayed with a different body. Change the key or send the original input.
429rate_limitedToo many requests. Back off and retry; do not tight-loop.
5xxinternalA server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice.

1. A tiny client

One helper that sends the token, unwraps data and raises on ok: false. The token comes from the token page (Copy token or Copy shell export); step 2 covers the kinds of token and minting one from code.

# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="stats-desk"
TOKEN="$SKILLSAFE_TOKEN"   # from https://stats-desk.skillsafe.ai/tokens.html

call() {                  # call <path> [json-body]
  if [ -n "$2" ]; then
    curl -sS -X POST "$BASE/$1" \
      -H "Authorization: Bearer $TOKEN" \
      -H "Content-Type: application/json" \
      -d "$2"
  else
    curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
  fi
}

2. Get a token

The easiest route is the token page: it shows the token this browser already holds, with Copy token and Copy shell export buttons, and a sign-in button for a personal token. A guest token, minted with POST /guest and {"slug":"stats-desk"}, can call /me and /estimate; the run is metered, so /run and /run-stream need a personal token.

# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
#   https://stats-desk.skillsafe.ai/tokens.html
#   export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
  -H "Content-Type: application/json" -d '{"slug":"stats-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}

3. Check the session and the balance

call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}

4. Price the run (free)

/estimate returns the model binding and the credits a run would reserve. It creates no job and charges nothing. Expect model_alias gpt-terra and markup_bps 1000 (a 10% markup). hold_credits is a reservation, not the price: it is held against your balance while the run executes and released afterwards. min_credits is the least balance that can start a run. What you actually pay is charged_credits, reported on the finished job and in the done event, and it is usually far lower than the hold. The body is the input object itself, with no {"input": …} wrapper. /estimate does not validate the body, so check the shape yourself: an object whose every value is a string, task equal to audit or draft, text and facts non-empty, and facts a JSON string that parses to an object (this is what the page's own guard, StatKit.mustBeObject, refuses to spend without).

# body.json is the input object itself - no {"input": ...} wrapper. Build it with
# make-body.js above, or by hand. estimate does not validate it, so check the shape first:
python3 -c 'import json;b=json.load(open("body.json"));assert isinstance(b,dict) and b.get("task") in ("audit","draft") and all(isinstance(v,str) for v in b.values()) and all(b.get(k,"").strip() for k in ("text","facts")) and isinstance(json.loads(b["facts"]),dict)'
INPUT=$(cat body.json)

call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
#   "markup_bps":1000,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
#   "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.

5. Run it, then poll

POST /run returns a job_id; poll GET /jobs/{id} until it is terminal. The reply is a string at data.output.output: JSON.parse it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the attempt number, stats-desk:<lane>:<hash>:a<attempt> (for example stats-desk:audit:lb7cqsc4tqiq:a1), so a retried request returns the same job instead of billing a second run. Use one key per distinct input: changed text, changed design notes, changed facts or a changed audit are a new hash, the same text in the other lane is a new key, and replaying an old key with a different body is a 409. The page uses StatKit.hashInput(body) for the hash (it covers task, target, journal, title, design, text, facts, audit and question; make-body.js prints the key); any stable digest of the body works from other languages. Leave retry_note out of the hash and bump the attempt instead.

# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])')   # audit or draft
KEY="stats-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"

JOB=$(curl -sS -X POST "$BASE/run" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')

while :; do
  OUT=$(call "jobs/$JOB")
  STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
  [ "$STATUS" = "succeeded" ] && break
  [ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
  sleep 2
done

# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
#   "output":{"output":"{\"lane\":\"audit\",\"verdict\":\"needs_revision\",\"headline\":\"...\", ...}"},
#   "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json

6. Or stream it

POST /run-stream takes the same body and headers and answers with server-sent events: job (the job id), delta (chunks of the reply) and done (the status, charged_credits, truncated and, when present, the full output). A browser page may receive only tick heartbeats and then done, never a delta, so take the reply from done.output.output when it is there, fall back to the concatenated deltas, and fall back again to GET /jobs/{id}.

# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
  -H "Authorization: Bearer $TOKEN" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $KEY" \
  -H "Accept: text/event-stream" \
  -d "$INPUT"

# event: job    {"job_id":"job_..."}
# event: delta  {"text":"{\"lane\":\"audit\",\"verdict\":\"needs_revision\",\"headline\":\"The rep"}
# event: done   {"status":"succeeded","charged_credits":...,"truncated":false}

7. Parse the reply

The reply is one JSON object, delivered as a string in data.output.output; you must JSON.parse it. The model is told to send no code fences, but tolerate them: strip a leading ```json and a trailing ```, keep everything from the first { to the last }, and parse that outer object. Then branch on lane: the envelope keys are the same for both lanes, the body keys are not.

# reply.json holds data.output.output from step 5. Strip any fence, keep the object:
python3 - <<'EOF'
import json, re
t = open("reply.json").read().strip()
t = re.sub(r"^```(?:json)?\s*", "", t, flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
r = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(r["lane"], r["verdict"], "-", r["headline"])
for i in r["issues"]:
    print(i["id"], i["priority"], i["ref"] or "-", "L%s" % i["line"], i["category"], "|", i["problem"])
if r["lane"] == "audit":
    print("UNIT", r["scope"]["unit"])
    for c in r["claims"]:
        print("CLAIM L%s" % c["line"], c["status"], c["claim"], "|", c["analysis"])
    open("revision.txt", "w").write(r["revision"]["text"])
else:
    open("statistical-analysis.txt", "w").write(r["draft_text"])
    print("UNRESOLVED", len(r["reporting_notes"]["unresolved"]))
    print("LEGEND", r["legend_template"])
for q in r["author_input_needed"]:
    print("ASK", q)
EOF

Invariants worth asserting

The web page holds every reply to the browser's facts and to your text before it shows it (recon.js: normalize(), then reconcile()). Do the same before you hold a manuscript on a verdict or paste the drafted text:

The page's checks are plain JavaScript, so the exact same reconciliation runs in Node:

// check-reply.js - hold a reply to the same checks the web page runs (recon.js reconcile()).
// Save recon.js and statkit.js next to this file.
// Usage: node check-reply.js body.json reply.json      (reply.json = data.output.output)
const fs = require("fs");
const K = require("./statkit.js");
const R = require("./recon.js");

const body = JSON.parse(fs.readFileSync(process.argv[2] || "body.json", "utf8"));
const facts = JSON.parse(body.facts);
// Rebuild the prescan from the numbered text the model saw: quotes are checked against these lines.
// (With clipped text, run this on the original text instead so line numbers stay exact.)
const text = body.text.split("\n").map((l) => l.replace(/^\d+\| /, "")).join("\n");
const A = K.analyze({ lane: body.task, target: body.target, journal: body.journal, title: body.title, design: body.design, text: text });
if (JSON.stringify(A.flags.map((f) => f.id)) !== JSON.stringify(facts.flags.map((f) => f.id)))
  console.warn("note: facts.flags differ from a fresh prescan of this text; the checks use the fresh one");

const reply = R.normalize(R.parseResult(fs.readFileSync(process.argv[3] || "reply.json", "utf8")), body.task);
const missing = R.SECTIONS[reply.lane].filter((k) => reply.present.indexOf(k) === -1);
if (missing.length) console.log("BAD  keys missing: " + missing.join(", "));
const rec = R.reconcile(reply, { analysis: A, input: body });
for (const i of rec.items) console.log((i.bad ? "BAD  " : i.review ? "LOOK " : "ok   ") + i.kind + ": " + i.text);
const look = rec.items.filter((i) => i.review).length;   // dismissals and placeholders: review by hand
console.log(reply.lane, reply.verdict, "-", rec.disagreements + missing.length, "disagreement(s),", look, "to review");
process.exit(rec.disagreements + missing.length ? 1 : 0);

Run on the worked examples below, both the audit and the draft replies report 0 disagreements, with one item to review each: the placeholders left for the authors to fill.

The output contract

Every key of the lane's contract is always present. Arrays may be empty ([], never a filler such as "None"). An enum is written "a|b|c": the reply carries exactly one of the values. Text is plain: no Markdown, no emoji, each string under 500 characters except revision.text and draft_text (under 3,500). The reply is only the JSON object, with no prose before or after it.

The common envelope (both lanes)

{"lane":"audit|draft",
 "verdict":"not_ready|needs_revision|ready",
 "headline":"...",
 "tldr":["...","..."],
 "issues":[{"id":"I1","ref":"S2","priority":"P0|P1|P2","category":"...","line":12,"evidence":"...","problem":"...","why":"...","fix":"..."}],
 ...the lane body...,
 "author_input_needed":["..."],
 "prescan_responses":[{"ref":"S1","verdict":"confirmed|dismissed","note":"..."}]}
keyshapewhat it holds
lanestringThe lane answered: your task, or the closest lane when task was missing or unknown.
verdictenumSee verdict and priority. In the draft lane it describes the drafted section.
headlinestringOne sentence on the state of the statistics.
tldrarray of strings2-5 bullets. When you sent a question, one bullet starts "Answer:".
issuesarray of objectsIds I1, I2, ... worst first. ref: the flag ids the issue answers, comma-separated, or "" for one the browser missed. category: snake_case, the flag's category when it answers one. line and evidence: one input line and a verbatim part of it (without the 12| prefix), or 0 and "" for an absence. why: what a statistical reviewer would challenge. Any p value discussed comes from facts.stats, cited by id. In the draft lane: only what the draft cannot resolve on its own.
author_input_neededarray of stringsShort factual questions, one per missing fact, each naming where the answer goes ("Which software and version ran the mixed model? (Statistical analysis)"). Every AUTHOR_INPUT_NEEDED placeholder is backed by one.
prescan_responsesarray of {ref, verdict, note}Exactly one per flag in facts.flags, in both lanes, with the reason and the line. Empty when no flags were sent.

The audit body

{"scope":{"input_reviewed":"...","boundary":"...","design":"...","unit":"..."},
 "claims":[{"line":6,"claim":"...","analysis":"...","n_unit":"...","status":"supported|underreported|overclaimed|not_assessable","note":"..."}],
 "revision":{"section":"statistical_analysis|results|legend|mixed","text":"...ready-to-paste text with AUTHOR_INPUT_NEEDED..."},
 "reviewer_risk":["..."]}
keywhat it holds
scopeinput_reviewed: what was supplied. boundary: what could not be assessed and why. design: the study design as the text states it (groups, treatments, time points, endpoints, paired or repeated structure). unit: the independent experimental unit and how biological replicates, technical replicates and subsamples relate to n, or "not stated".
claimsAt most 12, one per central result claim: the claim in a few words, the test or model reported for it ("not stated" if none), the n and its unit as reported, and a status: supported (analysis, n and uncertainty reported and matching the claim), underreported (may hold, key reporting missing), overclaimed (wording beyond the evidence) or not_assessable.
revisionReady-to-paste revised text for the part that most needs it. It keeps every supplied fact, removes overclaiming, and writes AUTHOR_INPUT_NEEDED for every missing fact; no number, test, correction or software appears that is not already in the text or design notes.
reviewer_risk1-4 sentences on what a statistical reviewer may still challenge after the revision.

The draft body

{"draft_text":"Statistical analysis. ...one string, AUTHOR_INPUT_NEEDED wherever a fact is missing...",
 "reporting_notes":{"n_definition":"...","tests_models":"...","multiple_comparisons":"...","software":"...","tails":"...","unresolved":["..."]},
 "legend_template":"...one sentence of figure-legend statistics text, with placeholders..."}
keywhat it holds
draft_textThe subsection as one string, conservative, in the order software, summary convention, sample size and replication, tests or models, multiplicity, exact reporting and thresholds, exclusions. A correction or model the audit recommends but the authors have not confirmed is an AUTHOR_INPUT_NEEDED, not a statement.
reporting_notesOne short string each for n_definition, tests_models, multiple_comparisons, software and tails, saying what the draft states and where it came from (line numbers or "design notes"); unresolved lists each placeholder left in the draft.
legend_templateOne sentence of figure-legend statistics (what points and bars show, n and its unit, test, correction, exact p values) in the same conservative style.

Verdict and priority

valuemeaning
not_readyAudit: any issue is P0. Draft: the independent unit, or the test for a central claim, is unknown (any open P0).
needs_revisionAudit: the worst issue is P1. Draft: placeholders remain only for P1 or P2 facts.
readyAudit: nothing worse than P2. Draft: no placeholders and no open issue.
P0 issueMust fix before submission: the independent unit wrong or undefined for a central claim, a paired, repeated or nested structure ignored, many comparisons with no correction or family, an interaction inferred from a difference in significance, undisclosed exclusions that could change the result, a reported p inconsistent with its statistic on the other side of 0.05, or an analysis that cannot be understood from the text.
P1 issueImportant: exact n missing, error bars or summary convention undefined, no effect size or uncertainty for key results, threshold-only p values, software missing, a p inconsistent with its statistic on the same side of 0.05, a small-sample limitation unacknowledged, trend language.
P2 issueClarity: inconsistent replicate terminology, panel-specific n missing, unclear ns labels, scattered statistical text.

The model may be stricter than facts.browser_verdict, never looser, unless it dismissed the flags that set it. The priority of an issue may also be stricter than its flag's (significance_difference is a medium flag, but an interaction inferred from a difference in significance is a P0).

Enums

wherevaluesthe page's fallback
laneaudit, draftthe lane asked
verdictnot_ready, needs_revision, readyderived from the worst issue
issues[].priorityP0, P1, P2P1
prescan_responses[].verdictconfirmed, dismissedconfirmed
claims[].statussupported, underreported, overclaimed, not_assessablenot_assessable
revision.sectionstatistical_analysis, results, legend, mixedmixed

The page also drops an issue with no problem, fix or evidence, a claim with no claim, and a flag response with no ref, and reads a line such as "L12" as 12 (anything unreadable as 0).

Worked example: audit

The page's "Psychology experiment, APA style" example (title Feedback timing and workload, target general, no design notes): a Statistical analysis paragraph and a Results section for a three-condition feedback study. The prescan recomputed five test results. Four check out (T2 to T5: two t-tests, a chi-square and a correlation), but T1, F(2, 87) = 3.12, p = .012, implies p = 0.0491 (range 0.0489 to 0.0494), so it is inconsistent, on the same side of 0.05. It raised three flags: S1 stat_inconsistent (medium), S2 trend_language (medium, line 11) and S3 randomization_blinding (low), so browser_verdict is needs_revision. There is no multiple_comparisons flag because the text names Bonferroni. This request is complete and sendable as make-body.js builds it; its Idempotency-Key from the page is stats-desk:audit:lb7cqsc4tqiq:a1.

The request body, with text and facts abridged:

{
 "task": "audit",
 "target": "general",
 "journal": "",
 "title": "Feedback timing and workload",
 "design": "",
 "text": "1| Statistical analysis\n2| Analyses were conducted in R 4.4.1. All tests were two-tailed with alpha = .05. Continuous ou... [abridged here: all 11 lines are shown decoded below]",
 "facts": "{\"target\":\"general\",\"browser_verdict\":\"needs_revision\",\"flags\":[{\"id\":\"S1\",\"severity\":\"medium\",\"priority\":\"P1\",\"category... [abridged here: decoded below]",
 "question": ""
}

Its text, decoded (the N| prefixes are part of the string):

1| Statistical analysis
2| Analyses were conducted in R 4.4.1. All tests were two-tailed with alpha = .05. Continuous outcomes are reported as mean ± SD. Pairwise comparisons after the omnibus ANOVA used Bonferroni correction.
3| 
4| Results
5| The final sample comprised 90 participants (n = 30 per condition); two participants who failed the attention check were excluded before analysis.
6| Perceived workload differed across the three feedback conditions, F(2, 87) = 3.12, p = .012, partial eta squared = .07.
7| Participants in the delayed-feedback condition reported higher workload than those in the immediate condition (M = 5.1, SD = 1.2 vs M = 4.4, SD = 1.3), t(58) = 2.17, p = .034, d = 0.56.
8| Workload did not differ between the immediate and no-feedback conditions, t(58) = 0.82, p = .42.
9| Condition was associated with task abandonment, chi2(2, N = 90) = 6.35, p = .042.
10| Self-reported effort correlated with workload, r(88) = .31, p = .003.
11| Accuracy was marginally significant in the delayed condition (p = .07), suggesting that delayed feedback drives poorer performance.

Its facts, decoded:

{
 "target": "general",
 "browser_verdict": "needs_revision",
 "flags": [
  {
   "id": "S1",
   "severity": "medium",
   "priority": "P1",
   "category": "stat_inconsistent",
   "message": "T1 (line 6): F(2, 87) = 3.12, p = .012 implies p = 0.0491 (range 0.0489 to 0.0494 allowing for rounding), not p = .012. Check the statistic, the df and the p value.",
   "lines": [
    6
   ]
  },
  {
   "id": "S2",
   "severity": "medium",
   "priority": "P1",
   "category": "trend_language",
   "message": "Trend language (\"marginally significant\", \"a trend toward significance\") recasts a non-significant result as a weak positive. Report the estimate and interval instead.",
   "lines": [
    11
   ]
  },
  {
   "id": "S3",
   "severity": "low",
   "priority": "P2",
   "category": "randomization_blinding",
   "message": "Animals or participants are compared, and the text says nothing about randomization or blinding.",
   "lines": []
  }
 ],
 "stats": [
  {
   "id": "T1",
   "line": 6,
   "test": "F",
   "df1": 2,
   "df2": 87,
   "value": 3.12,
   "reported": "F(2, 87) = 3.12, p = .012",
   "p_reported": "p = .012",
   "p_computed": 0.04913,
   "p_range": [
    0.04891,
    0.04936
   ],
   "status": "inconsistent"
  },
  {
   "id": "T2",
   "line": 7,
   "test": "t",
   "df1": 58,
   "df2": null,
   "value": 2.17,
   "reported": "t(58) = 2.17, p = .034",
   "p_reported": "p = .034",
   "p_computed": 0.03412,
   "p_range": [
    0.03372,
    0.03451
   ],
   "status": "consistent"
  },
  {
   "id": "T3",
   "line": 8,
   "test": "t",
   "df1": 58,
   "df2": null,
   "value": 0.82,
   "reported": "t(58) = 0.82, p = .42",
   "p_reported": "p = .42",
   "p_computed": 0.4156,
   "p_range": [
    0.4128,
    0.4184
   ],
   "status": "consistent"
  },
  {
   "id": "T4",
   "line": 9,
   "test": "chi2",
   "df1": 2,
   "df2": null,
   "value": 6.35,
   "reported": "chi2(2, N = 90) = 6.35, p = .042",
   "p_reported": "p = .042",
   "p_computed": 0.04179,
   "p_range": [
    0.04169,
    0.0419
   ],
   "status": "consistent"
  },
  {
   "id": "T5",
   "line": 10,
   "test": "r",
   "df1": 88,
   "df2": null,
   "value": 0.31,
   "reported": "r(88) = .31, p = .003",
   "p_reported": "p = .003",
   "p_computed": 0.002945,
   "p_range": [
    0.002495,
    0.003466
   ],
   "status": "consistent"
  }
 ],
 "n_mentions": [
  {
   "line": 5,
   "text": "n = 30 per condition); two particip",
   "unit": "unclear",
   "unit_word": "",
   "range": false
  },
  {
   "line": 9,
   "text": "N = 90) = 6.35, p = .042.",
   "unit": "unclear",
   "unit_word": "",
   "range": false
  }
 ],
 "tests_named": [
  "ANOVA"
 ],
 "corrections_named": [
  "Bonferroni"
 ],
 "software_named": [
  "R"
 ],
 "summary_conventions": [
  "mean ± s.d."
 ],
 "named_in_design_notes": {
  "tests": [],
  "corrections": [],
  "software": []
 },
 "tails_stated": true,
 "one_tailed_declared": false,
 "p_counts": {
  "exact": 6,
  "threshold": 0,
  "ns": 0
 },
 "clipped": null
}

The reply (data.output.output, parsed; shown in full apart from one phrase in reviewer_risk abbreviated with "..."). The model confirmed all three flags and carried each in an issue. It added three issues the browser missed (ref: ""), the most useful being I2: line 2 says the pairwise comparisons were Bonferroni-corrected, yet p = .034 is the unadjusted value T2 implies, so it is not stated whether the difference survives the correction. The worst issue is P1, so the verdict is needs_revision, matching the browser's floor. The revision rewrites line 11 as not statistically significant, drops "drives", keeps every reported number, and leaves the F value and p value of line 6 as placeholders rather than guessing which is wrong. check-reply.js reports 0 disagreements and one item to review: the placeholders.

{
 "lane": "audit",
 "verdict": "needs_revision",
 "headline": "The reporting is mostly complete and the degrees of freedom match the stated sample, but the omnibus F test's p value does not match its statistic, the Bonferroni status of the pairwise p values is unclear, and the accuracy sentence uses trend and causal language without a reported test.",
 "tldr": [
  "T1 (line 6): F(2, 87) = 3.12 implies p of about 0.0491 (range 0.0489 to 0.0494), not the reported p = .012; the F value, df or p value needs rechecking.",
  "Line 2 says pairwise comparisons used Bonferroni correction, but the p = .034 at line 7 matches the unadjusted value implied by the statistic (T2), so it is not stated whether the delayed versus immediate difference survives the stated correction.",
  "Line 11 calls p = .07 marginally significant and infers that delayed feedback drives poorer performance, with no test, statistic or n reported for accuracy.",
  "The t-tests, the chi-square test and the correlation are not described in the Statistical analysis section, and effect sizes or intervals are missing for lines 8 and 9."
 ],
 "issues": [
  {
   "id": "I1",
   "ref": "S1",
   "priority": "P1",
   "category": "stat_inconsistent",
   "line": 6,
   "evidence": "F(2, 87) = 3.12, p = .012, partial eta squared = .07.",
   "problem": "The reported p = .012 is inconsistent with F(2, 87) = 3.12; the statistic implies p of about 0.0491 (T1, range 0.0489 to 0.0494 allowing for rounding).",
   "why": "A reviewer will ask which of F, df or p is wrong; both values fall below .05, but the implied p lies very close to the threshold, which weakens the omnibus result that licenses the pairwise tests.",
   "fix": "Recheck the ANOVA output in R and report the F value, df and exact p value exactly as produced."
  },
  {
   "id": "I2",
   "ref": "",
   "priority": "P1",
   "category": "multiple_comparisons",
   "line": 7,
   "evidence": "t(58) = 2.17, p = .034, d = 0.56.",
   "problem": "Line 2 states that pairwise comparisons used Bonferroni correction, yet p = .034 matches the unadjusted two-tailed p implied by t(58) = 2.17 (T2, p_computed 0.0341), and only two of the three possible pairwise comparisons are reported.",
   "why": "A reviewer will ask whether the reported p values are Bonferroni-adjusted, how many comparisons formed the family, and whether the delayed versus immediate difference remains significant after the stated correction; the delayed versus no-feedback comparison is also missing.",
   "fix": "State whether the pairwise p values are adjusted or unadjusted, define the family of comparisons, report adjusted p values if the correction was applied, and report the delayed versus no-feedback comparison."
  },
  {
   "id": "I3",
   "ref": "S2",
   "priority": "P1",
   "category": "trend_language",
   "line": 11,
   "evidence": "Accuracy was marginally significant in the delayed condition (p = .07), suggesting that delayed feedback drives poorer performance.",
   "problem": "A non-significant result (p = .07 against the stated alpha = .05) is described as marginally significant and used to infer that delayed feedback drives poorer performance; no test, statistic, comparison group, n or estimate is reported for accuracy.",
   "why": "A reviewer will challenge recasting a non-significant result as a weak positive and drawing a causal conclusion from it; the analysis behind the p value cannot be identified from the text.",
   "fix": "Report the accuracy test, statistic, df, comparison groups, estimate and exact p value, describe the result as not statistically significant, and remove the causal inference."
  },
  {
   "id": "I4",
   "ref": "",
   "priority": "P1",
   "category": "tests_not_described",
   "line": 2,
   "evidence": "Pairwise comparisons after the omnibus ANOVA used Bonferroni correction.",
   "problem": "The Statistical analysis section names only the omnibus ANOVA; the type of pairwise t-test, the chi-square test at line 9, the correlation coefficient at line 10 and the accuracy analysis at line 11 are not described.",
   "why": "A reviewer cannot tell whether the pairwise tests were independent-samples t-tests or post hoc contrasts using the pooled ANOVA error, which correlation coefficient was used, or how accuracy was analysed.",
   "fix": "Name each test in the Statistical analysis section, including the pairwise test type, the chi-square test for abandonment, the correlation coefficient type and the accuracy analysis."
  },
  {
   "id": "I5",
   "ref": "",
   "priority": "P1",
   "category": "effect_size_uncertainty",
   "line": 8,
   "evidence": "Workload did not differ between the immediate and no-feedback conditions, t(58) = 0.82, p = .42.",
   "problem": "No effect size or interval is given for the null comparison at line 8 or for the chi-square result at line 9, and no confidence intervals are reported for any key result; line 8 also states no difference from a non-significant test.",
   "why": "Without an estimate and interval, a non-significant result cannot be read as evidence of no difference, and the size of the abandonment association cannot be judged.",
   "fix": "Report effect sizes with confidence intervals for lines 8 and 9 (and ideally for all key results), and word line 8 as no statistically significant difference."
  },
  {
   "id": "I6",
   "ref": "S3",
   "priority": "P2",
   "category": "randomization_blinding",
   "line": 0,
   "evidence": "",
   "problem": "The text does not state how participants were assigned to the three feedback conditions.",
   "why": "A reviewer will ask whether assignment was random, since between-condition differences are interpreted as effects of feedback timing.",
   "fix": "State how participants were allocated to conditions, or state that allocation was not randomized."
  }
 ],
 "scope": {
  "input_reviewed": "A Statistical analysis paragraph (line 2) and a Results section (lines 5 to 11) for a three-condition feedback timing study; no design notes, figures or legends were supplied.",
  "boundary": "Raw data, figures and legends were not supplied, so ANOVA assumptions, the source of the T1 mismatch, whether pairwise p values were adjusted, and the accuracy analysis are not assessable from the supplied material.",
  "design": "Three feedback conditions (delayed, immediate, no feedback) with 30 participants each after two attention-check exclusions; outcomes are perceived workload, task abandonment, self-reported effort and accuracy. The text suggests a between-participants design (t(58) and F(2, 87) are consistent with separate groups of 30), but the design and assignment method are not stated explicitly.",
  "unit": "The participant appears to be the independent unit (n = 30 per condition, N = 90), consistent with the reported degrees of freedom; no technical replicates or subsamples are described."
 },
 "claims": [
  {
   "line": 6,
   "claim": "Workload differed across the three feedback conditions",
   "analysis": "omnibus ANOVA",
   "n_unit": "90 participants, 30 per condition",
   "status": "underreported",
   "note": "The reported p = .012 is inconsistent with F(2, 87) = 3.12 (T1)."
  },
  {
   "line": 7,
   "claim": "Delayed feedback produced higher workload than immediate feedback",
   "analysis": "t-test, type not stated; Bonferroni correction stated at line 2",
   "n_unit": "30 participants per condition",
   "status": "underreported",
   "note": "p = .034 matches the unadjusted value (T2); whether it survives the stated Bonferroni correction is not stated."
  },
  {
   "line": 8,
   "claim": "No workload difference between immediate and no-feedback conditions",
   "analysis": "t-test, type not stated",
   "n_unit": "30 participants per condition",
   "status": "underreported",
   "note": "No effect size or interval; a non-significant test is worded as no difference."
  },
  {
   "line": 9,
   "claim": "Condition was associated with task abandonment",
   "analysis": "chi-square test, not described in the Statistical analysis section",
   "n_unit": "N = 90 participants",
   "status": "underreported",
   "note": "Statistic and p are consistent (T4), but no effect size is reported."
  },
  {
   "line": 10,
   "claim": "Self-reported effort correlated with workload",
   "analysis": "correlation, coefficient type not stated",
   "n_unit": "90 participants implied by r(88)",
   "status": "underreported",
   "note": "Statistic and p are consistent (T5); the correlation type and interval are not reported."
  },
  {
   "line": 11,
   "claim": "Delayed feedback drives poorer accuracy",
   "analysis": "not stated",
   "n_unit": "not stated",
   "status": "overclaimed",
   "note": "p = .07 is not significant at the stated alpha, and the causal wording goes beyond the reported evidence."
  }
 ],
 "revision": {
  "section": "mixed",
  "text": "Statistical analysis\nAnalyses were conducted in R 4.4.1. All tests were two-tailed with alpha = .05. Continuous outcomes are reported as mean ± SD. Perceived workload was compared across the three feedback conditions with an omnibus ANOVA (AUTHOR_INPUT_NEEDED: ANOVA design). Pairwise comparisons after the omnibus ANOVA used AUTHOR_INPUT_NEEDED (pairwise test type) with Bonferroni correction across AUTHOR_INPUT_NEEDED (number of comparisons in the family); reported pairwise p values are AUTHOR_INPUT_NEEDED (Bonferroni-adjusted or unadjusted). The association between condition and task abandonment was tested with a chi-square test, and the association between self-reported effort and workload with AUTHOR_INPUT_NEEDED (correlation coefficient type). Accuracy was analysed with AUTHOR_INPUT_NEEDED. Participants were allocated to conditions by AUTHOR_INPUT_NEEDED.\n\nResults\nThe final sample comprised 90 participants (n = 30 per condition); two participants who failed the attention check were excluded before analysis. Perceived workload differed across the three feedback conditions, F(2, 87) = AUTHOR_INPUT_NEEDED, p = AUTHOR_INPUT_NEEDED, partial eta squared = .07. Participants in the delayed-feedback condition reported higher workload than those in the immediate condition (M = 5.1, SD = 1.2 vs M = 4.4, SD = 1.3), t(58) = 2.17, p = .034 (AUTHOR_INPUT_NEEDED: adjusted or unadjusted; Bonferroni-adjusted p AUTHOR_INPUT_NEEDED), d = 0.56, confidence interval AUTHOR_INPUT_NEEDED. Workload did not differ significantly between the immediate and no-feedback conditions, t(58) = 0.82, p = .42, effect size AUTHOR_INPUT_NEEDED. The delayed and no-feedback conditions AUTHOR_INPUT_NEEDED (comparison result). Condition was associated with task abandonment, chi2(2, N = 90) = 6.35, p = .042, effect size AUTHOR_INPUT_NEEDED. Self-reported effort correlated with workload, r(88) = .31, p = .003. Accuracy in the delayed condition did not differ significantly from AUTHOR_INPUT_NEEDED (comparison group), AUTHOR_INPUT_NEEDED (test statistic and df), p = .07."
 },
 "reviewer_risk": [
  "If the corrected F test p value lies near .05, a reviewer may question how much ... the omnibus result can bear for the pairwise follow-ups.",
  "If the line 7 p value is unadjusted, the delayed versus immediate difference may not remain significant after the stated Bonferroni correction, which would change the central pairwise claim.",
  "A reviewer may ask whether assumptions of the ANOVA and t-tests were checked and how the accuracy outcome was defined and measured."
 ],
 "author_input_needed": [
  "Which of F, df or p at line 6 is correct, and what are the exact values from the R output? (Results, line 6)",
  "Are the pairwise p values at lines 7 and 8 Bonferroni-adjusted or unadjusted, and how many comparisons formed the family? (Statistical analysis and Results)",
  "What type of test was used for the pairwise comparisons, for example independent-samples t-tests or post hoc contrasts? (Statistical analysis)",
  "What is the result of the delayed versus no-feedback comparison? (Results)",
  "Which correlation coefficient was used for effort and workload? (Statistical analysis)",
  "What test, comparison group, statistic and df produced the accuracy p = .07? (Statistical analysis and Results, line 11)",
  "What are the effect sizes and confidence intervals for lines 7, 8 and 9? (Results)",
  "How were participants allocated to the three feedback conditions? (Statistical analysis or Methods)",
  "Was the ANOVA a one-way between-participants design? (Statistical analysis)"
 ],
 "prescan_responses": [
  {
   "ref": "S1",
   "verdict": "confirmed",
   "note": "Line 6 reports F(2, 87) = 3.12, p = .012, but T1 gives an implied p of about 0.0491 (range 0.0489 to 0.0494); both are below .05, so this is a same-side inconsistency (I1)."
  },
  {
   "ref": "S2",
   "verdict": "confirmed",
   "note": "Line 11 calls p = .07 marginally significant and infers that delayed feedback drives poorer performance (I3)."
  },
  {
   "ref": "S3",
   "verdict": "confirmed",
   "note": "Participants are compared across three conditions (line 5) and the text never states how they were allocated; blinding is less central for a self-report study, but allocation should be stated (I6)."
  }
 ]
}

Worked example: draft

The page's "Draft the Statistical analysis for the neuroscience paper" example (title Kcnq5 spine density (Fig. 2-3), target nature). It is the neuroscience Results section from the page's first example, plus three lines of design notes from the authors (3 mice per genotype, 18 cells from the same mice, unpaired t-tests in GraphPad Prism, nothing corrected, why two mice were excluded), plus the audit of that section, turned into text by Recon.auditText exactly as the "Draft the section from this audit" button does. The prescan found one test result, T1, t(4) = 2.10, P = 0.03, which implies p = 0.1037 (range 0.1031 to 0.1042): a decision_error. It raised thirteen flags, 3 high (stat_decision_error, n_subsample_unit for spines and cells, multiple_comparisons), 6 medium and 4 low, so browser_verdict is not_ready. Because the target is Nature, tails_unstated, df_missing and n_range are medium and repeats_unstated appears. Its Idempotency-Key from the page is stats-desk:draft:1cehxnmtrgdzo:a1.

The request body, with text, facts and audit abridged:

{
 "task": "draft",
 "target": "nature",
 "journal": "",
 "title": "Kcnq5 spine density (Fig. 2-3)",
 "design": "Kcnq5 knockout and wild-type littermates, 3 mice per genotype for imaging (one hippocampal slice per mouse), 3 mice per genotype for behaviour.\nSpines were measured on CA1 pyramidal neurons, several dendrites per mouse. mEPSCs: 18 cells from the same 3 mice per genotype.\nAll tests were unpaired t-tests in GraphPad Prism; nothing was corrected for multiple comparisons. The two excluded mice froze during the test phase.",
 "text": "1| Statistical analysis\n2| Data are presented as mean ± error bars. Statistical significance was assessed using unpaired... [abridged here: all 15 lines are shown decoded below]",
 "facts": "{\"target\":\"nature\",\"browser_verdict\":\"not_ready\",\"flags\":[{\"id\":\"S1\",\"severity\":\"high\",\"priority\":\"P0\",\"category\":\"stat_... [abridged here: decoded below]",
 "question": "",
 "audit": "I1 [S1] P0 stat_decision_error at line 10: The reported P = 0.03 does not match the statistic: per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.... [abridged here: decoded below]"
}

Its text, decoded:

1| Statistical analysis
2| Data are presented as mean ± error bars. Statistical significance was assessed using unpaired t-tests or one-way ANOVA in GraphPad Prism. P < 0.05 was considered significant. *P < 0.05, **P < 0.01, ***P < 0.001.
3| 
4| Results
5| Loss of Kcnq5 increased dendritic spine density in CA1 pyramidal neurons (n = 142 spines from 3 mice per genotype; P < 0.001, t-test; Fig. 2b).
6| Spine head volume was also larger in knockout neurons (n = 96 spines, P < 0.01; Fig. 2c), whereas spine length was not significantly different (ns; Fig. 2d).
7| Miniature EPSC frequency was higher in knockout slices (t = 3.41, P < 0.01; n = 18 cells), but mEPSC amplitude did not change (P > 0.05).
8| The increase in spine density was significant in males but not in females, indicating that the effect of Kcnq5 loss is sex-specific (Fig. 2e).
9| Across the five hippocampal subfields, knockout mice showed higher c-Fos counts in CA1, CA3 and DG (all P < 0.05) but not in CA2 or subiculum (Fig. 3a).
10| Novel object recognition was impaired in knockout mice (discrimination index 0.12 ± 0.05 vs 0.31 ± 0.04; t(4) = 2.10, P = 0.03; n = 3 mice per group).
11| Two outlier mice were excluded from the behavioural analysis.
12| Representative images are shown in Fig. 3b. These results prove that Kcnq5 controls spine formation.
13| 
14| Figure 2 legend
15| b-d, Spine density, head volume and length; each dot is one spine. Error bars, mean ± s.e.m. n = 3-4 mice per genotype.

Its audit, decoded (abbreviated here to its first four lines; the full string is 6,417 characters in 30 lines):

I1 [S1] P0 stat_decision_error at line 10: The reported P = 0.03 does not match the statistic: per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.1042, two-tailed), which is above 0.05. Fix: Recheck the statistic, df and P value against the Prism output and report the correct values; if the test was one-tailed this must be stated and justified in advance, otherwise revise the claim on line 10.
I2 [S2] P0 n_subsample_unit at line 5: Spines (lines 5-6) and cells (line 7) are used as n for t tests, while the independent unit is the mouse (3 per genotype); nothing states averaging per animal or a nested model. Fix: Analyse per-mouse means (n = mice) or use a mixed model with mouse as a random effect, and report the number of mice, cells and spines per group for each panel.
I3 [S4] P0 significance_difference at line 8: A sex-specific effect is inferred from the effect being significant in one sex and not the other; no interaction test is reported. Line 9 similarly contrasts subfields by significance. Fix: Report a test of the genotype by sex interaction (or a direct comparison of the effects) with n per sex, or remove the sex-specific interpretation.
I4 [S3] P0 multiple_comparisons at line 9: Five subfields are compared and reported with threshold P values only, with no correction or declared family of planned comparisons; no multiple-comparison correction is named anywhere in the text. Fix: Name the test for line 9, define the family of comparisons and state the correction applied (or that none was applied and why), and give exact P values per subfield.
[... abridged here: 26 more lines, the remaining issues I5-I16 and the author questions ...]

Its facts, decoded:

{
 "target": "nature",
 "browser_verdict": "not_ready",
 "flags": [
  {
   "id": "S1",
   "severity": "high",
   "priority": "P0",
   "category": "stat_decision_error",
   "message": "T1 (line 10): t(4) = 2.10, P = 0.03 implies p = 0.104 (two-tailed where applicable; range 0.103 to 0.104 allowing for rounding), so the reported p = 0.03 is on the wrong side of 0.05.",
   "lines": [
    10
   ]
  },
  {
   "id": "S2",
   "severity": "high",
   "priority": "P0",
   "category": "n_subsample_unit",
   "message": "n counts spines, cells and nothing says they were averaged per independent unit or modelled as nested - possible pseudoreplication, which shrinks p values.",
   "lines": [
    5,
    6,
    7
   ]
  },
  {
   "id": "S3",
   "severity": "high",
   "priority": "P0",
   "category": "multiple_comparisons",
   "message": "6 reported p value(s) or test result(s), but no multiple-comparison correction or declared family of planned comparisons is named.",
   "lines": [
    10
   ]
  },
  {
   "id": "S4",
   "severity": "medium",
   "priority": "P1",
   "category": "significance_difference",
   "message": "One effect is significant and another is not, and the text reads that as the two effects differing. A difference in significance is not a significant difference - test the interaction or compare the effects directly.",
   "lines": [
    6,
    7,
    8,
    9
   ]
  },
  {
   "id": "S5",
   "severity": "medium",
   "priority": "P1",
   "category": "tails_unstated",
   "message": "The text never says whether tests were one- or two-tailed - Nature's statistics section must state it.",
   "lines": []
  },
  {
   "id": "S6",
   "severity": "medium",
   "priority": "P1",
   "category": "df_missing",
   "message": "1 t or F statistic(s) printed without degrees of freedom; Nature asks for t and F values with their degrees of freedom.",
   "lines": [
    7
   ]
  },
  {
   "id": "S7",
   "severity": "medium",
   "priority": "P1",
   "category": "n_range",
   "message": "n is given as a range rather than exact values. Give the exact n for each group or panel, as Nature requires.",
   "lines": [
    15
   ]
  },
  {
   "id": "S8",
   "severity": "medium",
   "priority": "P1",
   "category": "exclusions_vague",
   "message": "Exclusions or outliers are mentioned without the rule, when it was set, or how many were removed.",
   "lines": [
    11
   ]
  },
  {
   "id": "S9",
   "severity": "medium",
   "priority": "P1",
   "category": "repeats_unstated",
   "message": "Representative results are shown without saying how many times the experiment was repeated - Nature asks for the repeat count.",
   "lines": [
    12
   ]
  },
  {
   "id": "S10",
   "severity": "low",
   "priority": "P2",
   "category": "ns_without_p",
   "message": "A result is called not significant (or ns) without its p value or effect estimate. Absence of significance is not evidence of no effect.",
   "lines": [
    6
   ]
  },
  {
   "id": "S11",
   "severity": "low",
   "priority": "P2",
   "category": "overclaim",
   "message": "Strength words (\"highly significant\", \"proves\", \"conclusively\") lean on the p value rather than the effect size and its uncertainty.",
   "lines": [
    12
   ]
  },
  {
   "id": "S12",
   "severity": "low",
   "priority": "P2",
   "category": "sem_used",
   "message": "s.e.m. describes the precision of a mean, not the spread of the data; with small n it makes groups look tighter. Pair it with n or show s.d. or a CI.",
   "lines": [
    15
   ]
  },
  {
   "id": "S13",
   "severity": "low",
   "priority": "P2",
   "category": "randomization_blinding",
   "message": "Animals or participants are compared, and the text says nothing about randomization or blinding.",
   "lines": []
  }
 ],
 "stats": [
  {
   "id": "T1",
   "line": 10,
   "test": "t",
   "df1": 4,
   "df2": null,
   "value": 2.1,
   "reported": "t(4) = 2.10, P = 0.03",
   "p_reported": "p = 0.03",
   "p_computed": 0.1037,
   "p_range": [
    0.1031,
    0.1042
   ],
   "status": "decision_error"
  }
 ],
 "n_mentions": [
  {
   "line": 5,
   "text": "n = 142 spines from 3 mice per genotype",
   "unit": "subsample",
   "unit_word": "spines",
   "range": false
  },
  {
   "line": 6,
   "text": "n = 96 spines, P < 0.01; Fig. 2",
   "unit": "subsample",
   "unit_word": "spines",
   "range": false
  },
  {
   "line": 7,
   "text": "n = 18 cells), but mEPSC amplit",
   "unit": "subsample",
   "unit_word": "cells",
   "range": false
  },
  {
   "line": 10,
   "text": "n = 3 mice per group).",
   "unit": "independent",
   "unit_word": "mice",
   "range": false
  },
  {
   "line": 15,
   "text": "n = 3-4 mice per genotype.",
   "unit": "independent",
   "unit_word": "mice",
   "range": true
  }
 ],
 "tests_named": [
  "t-test",
  "one-way ANOVA",
  "ANOVA"
 ],
 "corrections_named": [],
 "software_named": [
  "GraphPad Prism"
 ],
 "summary_conventions": [
  "mean ± s.e.m."
 ],
 "named_in_design_notes": {
  "tests": [
   "t-test"
  ],
  "corrections": [],
  "software": [
   "GraphPad Prism"
  ]
 },
 "tails_stated": false,
 "one_tailed_declared": false,
 "p_counts": {
  "exact": 1,
  "threshold": 5,
  "ns": 2
 },
 "clipped": null
}

The reply (parsed; shown in full). The drafted section keeps every fact the authors gave: the mouse as the independent unit, 3 mice per genotype, the 142 and 96 spines and 18 cells as reported, unpaired t-tests in GraphPad Prism, no correction, the freezing exclusion. Everything they did not give is an AUTHOR_INPUT_NEEDED: the Prism version, the tails, how spines and cells were analysed relative to the mouse, where one-way ANOVA was used. The audit recommended per-mouse means or a mixed model; the draft does not claim either was done. Four P0 issues remain open (the decision error on line 10, spines and cells as n, the sex-specific claim, the uncorrected subfields), so the drafted section is not_ready. All thirteen flags are answered. check-reply.js reports 0 disagreements (every number and method in the draft and the legend template comes from the text or the design notes) and one item to review: 20 placeholders against 15 author questions.

{
 "lane": "draft",
 "verdict": "not_ready",
 "headline": "The drafted section records the supplied design, software and the absence of multiplicity correction, but the central spine and mEPSC results were analysed with spines and cells as n, and the novel object recognition P value is inconsistent with its statistic, so it cannot yet be finalised.",
 "tldr": [
  "The draft states the mouse as the independent unit, 3 mice per genotype, unpaired t-tests in GraphPad Prism, no multiple-comparison correction and the freezing-based exclusion, all from the design notes.",
  "The spine (lines 5-6) and mEPSC (line 7) results use spines and cells as n; how they relate to the 3 mice per genotype must be resolved before the section is usable.",
  "Per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.1042), not the reported P = 0.03 on line 10; the value must be rechecked against the Prism output.",
  "The sex-specific claim (line 8) rests on a difference in significance and the five-subfield c-Fos comparison (line 9) has no correction or family definition.",
  "Nature asks for the tails of each test, exact n values, exact P values, t values with degrees of freedom and the number of repeats for representative images; these remain placeholders."
 ],
 "issues": [
  {
   "id": "I1",
   "ref": "S1",
   "priority": "P0",
   "category": "stat_decision_error",
   "line": 10,
   "evidence": "t(4) = 2.10, P = 0.03",
   "problem": "Per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.1042, two-tailed), on the other side of 0.05 from the reported P = 0.03.",
   "why": "A reviewer will question whether the novel object recognition impairment is supported and whether an undeclared one-tailed test was used.",
   "fix": "Recheck the statistic, df and P value against the Prism output and report the correct values; if the claim no longer holds at P < 0.05, revise the wording on line 10."
  },
  {
   "id": "I2",
   "ref": "S2",
   "priority": "P0",
   "category": "n_subsample_unit",
   "line": 5,
   "evidence": "n = 142 spines from 3 mice per genotype; P < 0.001, t-test",
   "problem": "Spines (lines 5-6) and cells (line 7) are used as n for t-tests, while the design notes give 3 mice per genotype as the source of all spines and all 18 cells.",
   "why": "Treating subsamples as independent inflates n and shrinks P values; with 3 mice per genotype the central claim may not be supported.",
   "fix": "State whether the analysis used per-mouse means (n = mice) or a model accounting for mouse; if neither, reanalyse and report mice, cells and spines per group for each panel."
  },
  {
   "id": "I3",
   "ref": "S4",
   "priority": "P0",
   "category": "significance_difference",
   "line": 8,
   "evidence": "The increase in spine density was significant in males but not in females, indicating that the effect of Kcnq5 loss is sex-specific",
   "problem": "A sex-specific effect is inferred from significance in one sex and not the other; no interaction test or n per sex is given.",
   "why": "A difference in significance is not a significant difference; a reviewer will ask for the genotype by sex interaction.",
   "fix": "Report a test of the genotype by sex interaction with n per sex, or remove the sex-specific interpretation."
  },
  {
   "id": "I4",
   "ref": "S3",
   "priority": "P0",
   "category": "multiple_comparisons",
   "line": 9,
   "evidence": "knockout mice showed higher c-Fos counts in CA1, CA3 and DG (all P < 0.05) but not in CA2 or subiculum",
   "problem": "Five subfields are compared with threshold P values; the design notes confirm no correction was applied and no family of planned comparisons is defined.",
   "why": "Uncorrected multiple comparisons raise the chance of false positives, and the subfield pattern is read from differences in significance.",
   "fix": "Define the family of comparisons and either apply a correction or state why none was applied, and report exact P values per subfield."
  },
  {
   "id": "I5",
   "ref": "",
   "priority": "P1",
   "category": "test_conflict",
   "line": 2,
   "evidence": "Statistical significance was assessed using unpaired t-tests or one-way ANOVA in GraphPad Prism.",
   "problem": "Line 2 names one-way ANOVA, while the design notes say all tests were unpaired t-tests.",
   "why": "A reviewer cannot tell which comparisons used which test, and Nature asks for F values and degrees of freedom for ANOVAs.",
   "fix": "Confirm whether one-way ANOVA was used; if so name the comparison and report F(df1, df2), otherwise remove it from line 2."
  },
  {
   "id": "I6",
   "ref": "S8",
   "priority": "P1",
   "category": "exclusions_vague",
   "line": 11,
   "evidence": "Two outlier mice were excluded from the behavioural analysis.",
   "problem": "The design notes give the reason (freezing during the test phase), but not whether the rule was prespecified, which genotype the mice came from, or the result with them included.",
   "why": "Undisclosed timing of an exclusion rule in a 3-per-group comparison could change the result.",
   "fix": "State the prespecified criterion, the genotype of each excluded mouse, and whether conclusions hold when they are included; describe them as excluded for freezing rather than as outliers."
  },
  {
   "id": "I7",
   "ref": "S5",
   "priority": "P1",
   "category": "tails_unstated",
   "line": 0,
   "evidence": "",
   "problem": "Neither the text nor the design notes say whether tests were one- or two-tailed.",
   "why": "Nature asks for the Methods statistics section to state whether tests were one- or two-tailed.",
   "fix": "Add the tails for all tests to the Statistical analysis."
  },
  {
   "id": "I8",
   "ref": "S7",
   "priority": "P1",
   "category": "n_range",
   "line": 15,
   "evidence": "n = 3-4 mice per genotype.",
   "problem": "The legend gives a range, which conflicts with line 5 and the design notes (3 mice per genotype).",
   "why": "Nature asks for exact n values, individual values rather than a range if n varied.",
   "fix": "Give the exact number of mice per genotype for each of panels b-d and reconcile with line 5."
  },
  {
   "id": "I9",
   "ref": "S6",
   "priority": "P1",
   "category": "df_missing",
   "line": 7,
   "evidence": "t = 3.41, P < 0.01; n = 18 cells",
   "problem": "The t value has no degrees of freedom and the P value is a threshold.",
   "why": "Nature asks for t values and degrees of freedom for t-tests.",
   "fix": "Report t(df) with the exact P value and the cells and mice per genotype."
  },
  {
   "id": "I10",
   "ref": "",
   "priority": "P1",
   "category": "threshold_p_values",
   "line": 6,
   "evidence": "n = 96 spines, P < 0.01; Fig. 2c",
   "problem": "All results except line 10 give threshold P values or ns, with no test statistics, effect sizes or intervals.",
   "why": "Nature asks for exact P values for significant and non-significant results where relevant; threshold values hide effect size and uncertainty.",
   "fix": "Report group means with the defined dispersion, the test statistic with df, and the exact P value for each comparison."
  },
  {
   "id": "I11",
   "ref": "S9",
   "priority": "P1",
   "category": "repeats_unstated",
   "line": 12,
   "evidence": "Representative images are shown in Fig. 3b.",
   "problem": "The number of times the representative measurement was repeated is not given.",
   "why": "Nature asks for the number of times representative measurements were repeated.",
   "fix": "State in the Fig. 3 legend how many independent experiments or mice the representative images reflect."
  },
  {
   "id": "I12",
   "ref": "",
   "priority": "P1",
   "category": "summary_convention_undefined",
   "line": 10,
   "evidence": "discrimination index 0.12 ± 0.05 vs 0.31 ± 0.04",
   "problem": "Line 2 says mean ± error bars without defining them; only the Fig. 2 legend says s.e.m., and the ± values on line 10 are undefined.",
   "why": "Nature asks for error bars to be defined throughout the figures.",
   "fix": "State what the ± values and error bars represent for every panel and Results value."
  },
  {
   "id": "I13",
   "ref": "S13",
   "priority": "P2",
   "category": "randomization_blinding",
   "line": 0,
   "evidence": "",
   "problem": "Randomization and blinding are not mentioned in the text or design notes.",
   "why": "Unblinded spine counting and behavioural scoring are a common reviewer concern.",
   "fix": "State whether allocation, spine quantification, electrophysiology and behavioural scoring were randomized or blinded."
  }
 ],
 "draft_text": "Statistical analysis. All statistical tests were performed in GraphPad Prism (version AUTHOR_INPUT_NEEDED). Spine density, head volume and length (Fig. 2b-d) are presented as mean ± s.e.m.; the summary convention for other panels, including the discrimination index values, is AUTHOR_INPUT_NEEDED, as is the n over which each s.e.m. was computed. Kcnq5 knockout mice and wild-type littermates were compared, and the independent experimental unit was the mouse. For spine imaging, 3 mice per genotype were analysed (one hippocampal slice per mouse); spines were measured on several dendrites of CA1 pyramidal neurons per mouse (142 spines for spine density and 96 spines for head volume; spines per genotype and per mouse: AUTHOR_INPUT_NEEDED; exact mice per genotype for Fig. 2b-d: AUTHOR_INPUT_NEEDED). mEPSCs were recorded from 18 cells from the same 3 mice per genotype (cells per genotype and per mouse: AUTHOR_INPUT_NEEDED). Behaviour was tested in 3 mice per genotype (number of mice tested before exclusion: AUTHOR_INPUT_NEEDED). Spines and cells are subsamples within each mouse; they were analysed as AUTHOR_INPUT_NEEDED (per-mouse means with n = mice, or a model that accounts for mouse). Comparisons between genotypes used unpaired t-tests (one- or two-tailed: AUTHOR_INPUT_NEEDED). Use of one-way ANOVA, the comparison it was applied to, and its F value and degrees of freedom: AUTHOR_INPUT_NEEDED. Test used to compare the genotype effect between males and females (Fig. 2e), with n per sex: AUTHOR_INPUT_NEEDED. No correction for multiple comparisons was applied, including across the five hippocampal subfields for c-Fos counts (Fig. 3a); family of planned comparisons and rationale: AUTHOR_INPUT_NEEDED. P < 0.05 was considered significant. Test statistics are reported with degrees of freedom and exact P values for all comparisons, including non-significant ones (values: AUTHOR_INPUT_NEEDED). Two mice were excluded from the behavioural analysis because they froze during the test phase; whether this criterion was set before analysis, the genotype of each excluded mouse and the result with them included: AUTHOR_INPUT_NEEDED. Representative images (Fig. 3b) reflect AUTHOR_INPUT_NEEDED independent experiments. Randomization of allocation and blinding during spine quantification, electrophysiology and behavioural scoring: AUTHOR_INPUT_NEEDED.",
 "reporting_notes": {
  "n_definition": "Mouse as the independent unit; 3 mice per genotype for imaging with one slice per mouse, 18 mEPSC cells from the same mice, and 3 mice per genotype for behaviour, from the design notes; 142 and 96 spines from lines 5-6; per-mouse breakdown left as placeholders.",
  "tests_models": "Unpaired t-tests from the design notes and line 2; the one-way ANOVA on line 2 conflicts with the design notes and is left as a placeholder, as is the analysis of spines and cells relative to mouse and the sex comparison.",
  "multiple_comparisons": "States that no correction was applied, from the design notes; the family of planned comparisons and rationale are a placeholder.",
  "software": "GraphPad Prism from line 2 and the design notes; version is a placeholder.",
  "tails": "Not stated in text or design notes; left as a placeholder.",
  "unresolved": [
   "GraphPad Prism version",
   "Summary convention for panels other than Fig. 2b-d, including the ± values of the discrimination index, and the n over which s.e.m. was computed",
   "Spines per genotype and per mouse for spine density and head volume",
   "Exact number of mice per genotype for Fig. 2b-d (line 5 gives 3, the legend gives 3-4)",
   "mEPSC cells per genotype and per mouse",
   "Number of behaviour mice tested before exclusion",
   "How spines and cells were analysed relative to mouse (per-mouse means or a nested model)",
   "Whether tests were one- or two-tailed",
   "Whether and where one-way ANOVA was used, with F and degrees of freedom",
   "Test for the genotype by sex comparison and n per sex",
   "Family of planned comparisons and rationale for no correction",
   "Exact test statistics, degrees of freedom and P values for each comparison",
   "Whether the freezing exclusion was prespecified, genotype of each excluded mouse, and result with them included",
   "Number of independent experiments behind the representative images in Fig. 3b",
   "Randomization and blinding"
  ]
 },
 "legend_template": "Each dot represents one AUTHOR_INPUT_NEEDED (mouse, cell or spine); bars show mean ± s.e.m. of n = AUTHOR_INPUT_NEEDED mice per genotype (AUTHOR_INPUT_NEEDED spines or cells per mouse); AUTHOR_INPUT_NEEDED (one- or two-tailed) unpaired t-test, no correction for multiple comparisons; exact P values: AUTHOR_INPUT_NEEDED.",
 "author_input_needed": [
  "Which GraphPad Prism version was used? (Statistical analysis)",
  "Were spine and cell measurements averaged per mouse or analysed with a model accounting for mouse, and with what model? (Statistical analysis)",
  "How many spines per genotype and per mouse contributed to spine density and head volume? (Statistical analysis; Results, lines 5-6)",
  "What are the exact numbers of mice per genotype for Fig. 2b-d, given line 5 says 3 and the legend says 3-4? (Figure 2 legend)",
  "How many of the 18 mEPSC cells came from each genotype and each mouse, and what is the df for t = 3.41? (Statistical analysis; Results, line 7)",
  "Were tests one- or two-tailed? (Statistical analysis)",
  "Was one-way ANOVA used, for which comparison, and what are F and its degrees of freedom? (Statistical analysis; Results)",
  "What are the correct t statistic, df and exact P value for the novel object recognition comparison? (Results, line 10)",
  "Was a genotype by sex interaction tested, and what are the n per sex? (Statistical analysis; Results, line 8)",
  "What is the family of planned comparisons for the five c-Fos subfields and why was no correction applied? (Statistical analysis; Results, line 9)",
  "What are the exact P values and test statistics for all comparisons, including spine length and mEPSC amplitude? (Results)",
  "What do the ± values on line 10 and the error bars in panels other than Fig. 2b-d represent? (Statistical analysis; legends)",
  "How many behaviour mice were tested before exclusion, was the freezing criterion set before analysis, which genotype were the two excluded mice, and does the result change with them included? (Statistical analysis; Results, line 11)",
  "How many times were the representative measurements in Fig. 3b repeated? (Figure 3 legend)",
  "Were allocation, spine quantification, electrophysiology and behavioural scoring randomized or blinded? (Statistical analysis)"
 ],
 "prescan_responses": [
  {
   "ref": "S1",
   "verdict": "confirmed",
   "note": "Line 10 reports t(4) = 2.10, P = 0.03; per T1 the statistic implies p_computed 0.1037 (range 0.1031 to 0.1042), above 0.05."
  },
  {
   "ref": "S2",
   "verdict": "confirmed",
   "note": "Lines 5-7 use spines and cells as n; the design notes say all came from 3 mice per genotype and describe no per-mouse averaging or nested model."
  },
  {
   "ref": "S3",
   "verdict": "confirmed",
   "note": "Line 9 compares five subfields with threshold P values and the design notes confirm nothing was corrected for multiple comparisons; no family is defined."
  },
  {
   "ref": "S4",
   "verdict": "confirmed",
   "note": "Line 8 infers a sex-specific effect from significance in males but not females without an interaction test; line 9 contrasts subfields the same way."
  },
  {
   "ref": "S5",
   "verdict": "confirmed",
   "note": "Neither the text nor the design notes state whether tests were one- or two-tailed."
  },
  {
   "ref": "S6",
   "verdict": "confirmed",
   "note": "Line 7 gives t = 3.41 without degrees of freedom."
  },
  {
   "ref": "S7",
   "verdict": "confirmed",
   "note": "Line 15 gives n = 3-4 mice per genotype, a range that also conflicts with line 5 and the design notes (3 per genotype)."
  },
  {
   "ref": "S8",
   "verdict": "confirmed",
   "note": "Line 11 gives no rule; the design notes add that the mice froze during the test phase, but not whether the rule was prespecified or their genotype."
  },
  {
   "ref": "S9",
   "verdict": "confirmed",
   "note": "Line 12 shows representative images in Fig. 3b without the number of repeats."
  },
  {
   "ref": "S10",
   "verdict": "confirmed",
   "note": "Line 6 reports spine length as ns and line 7 mEPSC amplitude as P > 0.05, without exact P values or estimates."
  },
  {
   "ref": "S11",
   "verdict": "confirmed",
   "note": "Line 12 says these results prove that Kcnq5 controls spine formation, beyond density data from 3 mice per genotype."
  },
  {
   "ref": "S12",
   "verdict": "confirmed",
   "note": "Line 15 uses s.e.m. with each dot one spine and 3-4 mice per genotype, so the n behind the s.e.m. is unclear."
  },
  {
   "ref": "S13",
   "verdict": "confirmed",
   "note": "Mice are compared and neither the text nor the design notes mention randomization or blinding."
  }
 ]
}

Truncation and partial results

When the balance sits between min_credits and hold_credits, the run is not refused: it executes with a reduced output cap and reports truncated: true, in the done event of /run-stream and on the job from GET /jobs/{id}. What you hold is then a prefix of the reply, and it will not parse as it stands. The web page closes the cut-off JSON (Recon.closeJson in recon.js: close an open string, drop a dangling comma, give a dangling key null, close every open array and object, and if that still does not parse, cut back to the previous comma and try again), parses what is left, and shows the sections that arrived as "N of M sections recovered", out of the lane's 11 keys (10 for draft). A stream that ends early with an error gets the same treatment on the deltas received so far.

The keys arrive in contract order, so a cut usually costs the tail: prescan_responses and author_input_needed go first, then the lane body. A truncated audit can therefore look complete while flags are unanswered, and a cut inside revision.text or draft_text leaves you half a section with placeholders that no question backs. Never paste a truncated section: check the flag, top up, and resubmit with the attempt suffix on the Idempotency-Key incremented (stats-desk:audit:<hash>:a2).

If a complete reply will not parse as one JSON object, the page retries once, as the next attempt, with a retry_note saying what was wrong: "Your previous reply was not the single valid JSON object the instructions require (the parse error). Reply again with ONLY the JSON object for task 'lane' - no prose, no code fences; every key present (empty arrays where there is nothing to say)." Do the same: keep the hash, bump the attempt, add retry_note. It never retries a truncated reply that way; that one needs credits, not a reformat. If the retry still does not parse, the page shows the raw reply.