Drive Stats Desk from your own code
A reporting review, not a reanalysis. The model reads your text; it never sees raw
data, never runs a test and never computes a p value itself. Your manuscript travels in
text (up to 24,000 characters of numbered lines; longer text keeps the head and the tail
and drops the middle with a marker), so leave out anything you may not share, such as unpublished
participant details, before you send it. It is not medical, regulatory or clinical-trial advice.
Everything the web page does is available over HTTP. Send a Statistical analysis subsection, Results
paragraphs, figure legends or reviewer comments and get the same review back: a verdict
(not_ready, needs_revision, ready), a one-sentence headline, a
short TL;DR, issues ranked P0 to P2 that each cite a line, quote it and say
what a statistical reviewer would challenge and how to fix it, the lane's body (the scope, a claim
map, a ready-to-paste revision and the reviewer risk for an audit; the drafted section, reporting
notes and a legend template for a draft), the factual questions only the authors can answer, and one
answer per browser flag. The natural use is a pre-submission check in a writing pipeline: run the
statistics text through here and hold the manuscript on not_ready.
The model does not do the arithmetic. Every reported t, F, chi-square, r and z
statistic with its degrees of freedom is found first by the page's free prescan in
statkit.js, which turns it back into the p value it implies (allowing for the rounding
of both numbers, with a one-tailed allowance and a decision-error test at 0.05) and marks it
consistent, inconsistent, decision_error and so on. The same
prescan flags n counted as cells or spines, missing multiple-comparison corrections, undefined error
bars, threshold-only p values, trend language and more. The model is told to quote those computed
values by id (T1) and never to state a p value of its own. The result is sent as
facts, a JSON string. See the facts string.
Two lanes: the task field
Every request names its lane in task. Stats Desk has two, with two different reply bodies.
| task | what it is for | lane body in the reply |
|---|---|---|
audit | Review the statistical reporting of the supplied text: the independent unit and how n relates to it, paired or nested structure, multiplicity, reported p values against their statistics, effect sizes and uncertainty, error-bar definitions, exclusions, software, tails, overclaiming. Returns issues ranked P0/P1/P2, a map of the central claims, and ready-to-paste revised text for the part that most needs it, with AUTHOR_INPUT_NEEDED wherever a fact is missing. | scope, claims, revision, reviewer_risk. |
draft | Draft a ready-to-paste Statistical analysis subsection from the text, the authors' design notes (design) and, if given, an earlier audit (audit), in a fixed order: software; summary convention; sample size and replication; tests or models per comparison; multiplicity; exact reporting and thresholds; exclusions. A method the audit recommends but the authors have not confirmed is written as AUTHOR_INPUT_NEEDED, never as done. The verdict describes the DRAFTED section, and issues lists only what the draft cannot resolve on its own. | draft_text, reporting_notes, legend_template. |
Both lanes share one envelope (lane, verdict, headline,
tldr, issues, author_input_needed,
prescan_responses); see the output contract. If
task is missing or unrecognised, the model chooses the closest lane, answers with that
lane's contract (never a blend) and names it in lane. Do not rely on that: always send
"audit" or "draft". The page's own guard (StatKit.mustBeObject)
refuses any other value, and its checks flag a reply whose lane differs from the
task asked.
The prescan is the same for both lanes: facts depends only on the text, the target and
the design notes. The same text run in two lanes is still two different runs, with two different
Idempotency-Keys. The page's own flow is audit first, then draft: its "Draft the section
from this audit" button sends the audit's issues and author questions as plain text
(Recon.auditText, one line per issue and one per question) in audit, with the
same text and design notes.
Input fields
The body is one flat JSON object, built in the page by StatKit.buildInput. Every value is
a string. Required: task, text, facts.
| field | type | required | what it holds |
|---|---|---|---|
task | string | yes | "audit" or "draft"; see the lanes. |
target | string | no | "general", "nature" (the flagship journal Nature) or "other". Anything else is treated as general by the prescan. Only with nature does the model say "Nature asks for", and only for Nature's initial-submission statistical items (tests and tails named, error bars defined, exact n, repeat counts for representative results, exact P values, F and t with their degrees of freedom, replicates defined); for other targets it calls these good reporting practice. The prescan also raises the severity of tails_unstated, df_missing, test_stat_missing and n_range to medium for Nature, and adds repeats_unstated. |
journal | string | no | The journal's name when target is other, else "". Cut at 120 characters. |
title | string | no | A label for the manuscript ("Feedback timing and workload"). Context only. Cut at 200 characters. |
design | string | no | The authors' own notes on the design: groups, units, replicates, software, what was or was not corrected, why samples were excluded. These count as facts just like the text, so a number or method stated here may appear in the drafted text. Cut at 3,000 characters on a word, the cut marked [... the rest of the design notes (N characters) was cut ...]. Most useful in the draft lane. |
text | string | yes | The manuscript text with every line prefixed by its number and "| " ("12| Spine density was higher ..."), so each issue can cite a real line. At most 24,000 characters: longer text keeps about 60% of the budget from the head and the rest from the tail, on whole lines, with a marker line in between: [... lines A-B were not sent (N lines); the facts were computed from the whole text ...]. Trailing blank lines are dropped before numbering. |
facts | string | yes | A JSON string (the output of JSON.stringify), never an object. In the page it is the browser's free prescan of the WHOLE text, including any clipped middle; its keys are listed below. |
audit | string | draft only | The issues from an earlier audit, as plain text, one per line ("I1 [S1] P0 stat_decision_error at line 10: ... Fix: ...", then "Author input needed: ..." lines). Cut at 8,000 characters. May be "": the model then drafts from the text, the design notes and the flags. The page sends this key only in the draft lane. |
question | string | no | Your own question, answered inside tldr as a bullet starting "Answer:". Cut at 1,200 characters on a word; the page sends "" when empty. |
retry_note | string | no | Only on a reformat retry, after a reply that could not be parsed: say what was wrong. The page adds it to the body of the second attempt and never on a first run. See truncation. |
The facts string
In the web page, facts is computed by the browser before you pay for anything: the free
prescan (statkit.js) reads the whole text, recomputes every test result it can parse,
raises the flags you see on the page, lists what the text names, and serializes the result. An API
caller has two options: build the same object with statkit.js, which runs unchanged in
Node (see building the body), or send a minimal one yourself.
| key | what it holds |
|---|---|
target | The target as sent, with the journal in brackets for other ("other (eLife)"). |
browser_verdict | The prescan's hint: not_ready (any high flag), needs_revision (any medium), else ready. The model's verdict is never looser unless it dismissed the flags that set it. |
flags | Up to 24 flags, worst first: id (S1, S2, ...), severity (high, medium, low), priority (P0, P1, P2, one to one with severity), category, message and lines (up to 12 line numbers; [] for an absence). The categories are in the next table. |
stats | Every test result the browser found, as {id, line, test, df1, df2, value, reported, p_reported, p_computed, p_range, status}. id is T1, T2, ...; test is t, F, chi2, r or z; df1/df2 are null when absent; reported is the matched text (up to 140 characters); p_reported is "p = .034", "p < 0.01", "ns" or "none". p_computed is the p value the statistic implies (two-tailed for t, z and r) to four significant figures, and p_range is [low, high] allowing for the rounding of the printed statistic; both are null when status is impossible. These are the ONLY computed p values the model may quote. |
n_mentions | Up to 40 sample sizes found as n = ...: {line, text, unit, unit_word, range}. unit is independent (mice, patients, participants, cultures, independent experiments...), subsample (cells, spines, fields, wells, images, technical replicates...) or unclear; unit_word is the word that followed; range is true for n = 3-4. |
tests_named, corrections_named, software_named, summary_conventions | What the text names: tests ("t-test (paired)", "one-way ANOVA", "mixed-effects model", ...), multiple-comparison corrections ("Bonferroni", "Holm", "Benjamini-Hochberg / FDR", "Tukey", ...), software ("R", "GraphPad Prism", "SPSS", ...) and summary conventions ("mean ± s.d.", "mean ± s.e.m.", "confidence interval", "median and IQR", "box plot convention"). |
named_in_design_notes | {tests, corrections, software}: the same lists read from design. |
tails_stated, one_tailed_declared | Whether the text says one- or two-tailed (or sided) anywhere, and whether it says one-tailed. With a declared one-tailed test, a statistic that matches its p only as a one-tailed value counts as consistent. |
p_counts | {exact, threshold, ns}: how many p values are printed as p = x, as p < x, and as ns / not significant. |
clipped | null, or {total_lines, dropped} (dropped as "A-B") when the middle of the text was cut from text. |
stats[].status | meaning |
|---|---|
consistent | The reported p lies within what the statistic implies, allowing for rounding of both numbers (or as a one-tailed p when the text declares one-tailed tests). |
inconsistent | It does not, but both are on the same side of 0.05 (flag stat_inconsistent, medium). |
decision_error | It does not, and the reported p is on the other side of 0.05 from the implied one: a significant result that is not, or the reverse (flag stat_decision_error, high). |
one_tailed_only | A t, z or r result that matches only if the p is one-tailed, and the text never says one-tailed tests were used (flag stat_one_tailed_only, medium). |
impossible | The value cannot occur: r beyond ±1, a negative F or chi-square, a missing df where one is required, or p above 1 (flag stat_impossible, high). |
no_p | A statistic with no p value next to it; p_computed is still given. |
| flag category | severity | what the prescan matched |
|---|---|---|
stat_impossible | high | A stats entry with status impossible. |
stat_decision_error | high | A stats entry with status decision_error. |
stat_inconsistent | medium | A stats entry with status inconsistent. |
stat_one_tailed_only | medium | A stats entry with status one_tailed_only. |
n_subsample_unit | high, or low when the text names a mixed, nested or per-animal averaging approach | n counts cells, spines, fields, images or technical replicates (possible pseudoreplication). |
multiple_comparisons | high | Six or more p values or test results, or wording about many comparisons, with no correction or declared family named. |
p_zero | medium | A p value printed as zero (p = 0.000). |
p_threshold_only | medium | Every p value is a threshold; no exact values. |
trend_language | medium | "Marginally significant", "a trend toward significance". |
significance_difference | medium | One effect significant, another not, read as the effects differing. |
summary_undefined | medium | Error bars or ± values never defined as s.d., s.e.m. or a CI. |
exclusions_vague | medium | Exclusions or outliers without the rule, its timing or the count. |
stars_undefined | medium | Significance stars whose thresholds are never defined. |
n_missing | medium | Test results but no sample size anywhere. |
repeats_unstated | medium (Nature only) | Representative results with no repeat count. |
tails_unstated | medium for Nature, else low | t, z, r or rank tests with no statement of tails. |
df_missing | medium for Nature, else low | A t or F printed without degrees of freedom. |
test_stat_missing | medium for Nature, else low | p values for t-tests or ANOVAs with no t or F statistic at all. |
n_range | medium for Nature, else low | n given as a range (n = 3-4). |
ns_without_p | low | "ns" or "not significant" with no p value or estimate. |
overclaim | low | "Highly significant", "proves", "conclusively". |
association_as_cause | low | An association written in causal or mechanistic terms. |
randomization_blinding | low | Animals or participants compared with no word on randomization or blinding. |
sem_used | low | s.e.m. used as the summary convention. |
normality_small_n | low | Normality asserted or tested with n of 6 or fewer per group. |
software_missing | low | Inference reported with no statistical software named. |
The flags drive the reply. Every flag must come back exactly once in
prescan_responses (confirmed or dismissed, with the reason and
the line), and every confirmed high or medium flag must be carried by an
issue with that flag's id in ref. They are pattern matches and can be wrong: the model
is told to dismiss honestly, for example n counting cells when a mixed model with animal as a random
effect is described, or a one_tailed_only stat where the Methods do state one-tailed tests.
Sending facts without the prescan. Empty lists are allowed. A minimal
facts, as the JSON string you put in the body:
"facts": "{\"flags\":[],\"stats\":[]}"
It costs you the reconciliation the page does. With no flags there is nothing to confirm or dismiss,
prescan_responses comes back [], the verdict has no floor, and the issues
rest on the model's own reading of the text. With no stats, the model has no
browser-computed p value to cite, and it is told never to compute one, so it can say a p value looks
doubtful but cannot tell you what the statistic implies. Your own checks
lose their reference too. Running statkit.js is free and gets you all of it.
Building the body
The surest way to match the page is to run the page's own engine. Save statkit.js from
this site (it runs unchanged in Node via require()) and give analyze the
same fields the page's form has; buildInput then numbers and clips the text, cuts the
text fields and serializes the facts:
| set field | what it holds |
|---|---|
lane | audit or draft (anything else is analyzed as audit). Becomes task, and decides whether audit is sent. |
target, journal | As in the input fields. |
title | The manuscript label. |
design | The design notes, raw. |
text | The raw text, without line numbers. |
// make-body.js - build the run body with the SAME engine the web page uses.
// Save statkit.js from https://stats-desk.skillsafe.ai/ next to this file.
// Usage: node make-body.js survey.txt audit general
// node make-body.js mouse.txt draft nature design.txt audit.txt
const fs = require("fs");
const K = require("./statkit.js");
const [file = "survey.txt", lane = "audit", target = "general", designFile, auditFile] = process.argv.slice(2);
const A = K.analyze({
lane: lane, // audit or draft
target: target, // general, nature or other
journal: "", // the journal's name when target is other
title: "Feedback timing and workload",
design: designFile ? fs.readFileSync(designFile, "utf8") : "",
text: fs.readFileSync(file, "utf8") // raw text; buildInput numbers the lines and clips
});
const audit = lane === "draft" && auditFile ? fs.readFileSync(auditFile, "utf8") : "";
const body = K.buildInput(A, { question: "", audit: audit });
fs.writeFileSync("body.json", JSON.stringify(K.mustBeObject(body)));
console.log(A.flags.length, "flags,", A.stats.length, "stats; browser verdict", A.hint);
console.log("Idempotency-Key: stats-desk:" + body.task + ":" + K.hashInput(body) + ":a1");
Run on the texts from the page's examples, this produces the worked requests below, byte for byte:
the psychology experiment (audit, target
general, hash lb7cqsc4tqiq) and the
neuroscience draft (draft, target nature, with
design notes and an audit, hash 1cehxnmtrgdzo). Trailing blank lines are dropped before
numbering, so a trailing newline does not change the hash; any other edit does.
mustBeObject is the page's own guard: it throws unless the body is a plain object with
task audit or draft, non-empty text and a
facts string. From another language, send the same field names, number the lines
yourself ("1| ", "2| ", ...) and build facts with the keys
above, or the minimal one.
For the draft lane, turn an audit reply into the audit text the same way the page's
"Draft the section from this audit" button does:
// audit-text.js - turn an audit reply into the draft lane's audit text, as the page does.
// Save recon.js next to statkit.js. Usage: node audit-text.js reply.json
const fs = require("fs");
const R = require("./recon.js");
const reply = R.normalize(R.parseResult(fs.readFileSync(process.argv[2] || "reply.json", "utf8")), "audit");
fs.writeFileSync("audit.txt", R.auditText(reply));
console.log(reply.issues.length, "issues and", reply.author_input_needed.length, "questions written to audit.txt");
Base URL and the envelope
Every endpoint lives under https://api.skillsafe.ai/v1/app-api and every response uses
the same envelope, so one helper covers the whole API:
{"ok": true, "data": {"job_id": "job_...", "status": "queued"}}
{"ok": false, "error": {"code": "payment_required", "message": "..."}}
The token is minted for this app (the guest endpoint takes {"slug":"stats-desk"} in
its body), so no slug header is needed afterwards. Send it as Authorization: Bearer ….
The input object IS the request body. There is no {"input": …}
wrapper. A wrapped body is answered with an unknown field 'input' warning, and the
model never sees your text.
Error codes
| status | code | what to do |
|---|---|---|
| 400 | validation_error | A field is missing or the wrong type. Every field is a string: facts must be a JSON-encoded string, not an object. |
| 401 | unauthorized | The token is missing, malformed or expired. Get a new one from the token page. |
| 402 | payment_required | The balance is below min_credits. Call /estimate first and top up. |
| 403 | forbidden | The token is valid but not for this app, or a guest token tried a metered run. A guest cannot run; sign in for a personal token. |
| 404 | not_found | Unknown job id, or the app slug does not exist. |
| 409 | conflict | The same Idempotency-Key was replayed with a different body. Change the key or send the original input. |
| 429 | rate_limited | Too many requests. Back off and retry; do not tight-loop. |
| 5xx | internal | A server-side failure. Retry with the SAME Idempotency-Key so you are not billed twice. |
1. A tiny client
One helper that sends the token, unwraps data and raises on ok: false.
The token comes from the token page (Copy token or
Copy shell export); step 2 covers the kinds of token and minting one from code.
# Every call is the same three things: the base URL, your bearer token,
# and a JSON body. Keep the token in a shell variable.
BASE="https://api.skillsafe.ai/v1/app-api"
SLUG="stats-desk"
TOKEN="$SKILLSAFE_TOKEN" # from https://stats-desk.skillsafe.ai/tokens.html
call() { # call <path> [json-body]
if [ -n "$2" ]; then
curl -sS -X POST "$BASE/$1" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-d "$2"
else
curl -sS "$BASE/$1" -H "Authorization: Bearer $TOKEN"
fi
}
import json, os, urllib.error, urllib.request
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "stats-desk"
TOKEN = os.environ.get("SKILLSAFE_TOKEN", "YOUR_TOKEN") # from https://stats-desk.skillsafe.ai/tokens.html
def call(path, body=None):
"""Returns the unwrapped `data`, or raises with the API error code."""
data = json.dumps(body).encode() if body is not None else None
req = urllib.request.Request(f"{BASE}/{path}", data=data, method="POST" if body is not None else "GET")
req.add_header("Authorization", f"Bearer {TOKEN}")
if body is not None:
req.add_header("Content-Type", "application/json")
try:
with urllib.request.urlopen(req) as r:
payload = json.load(r)
except urllib.error.HTTPError as e:
payload = json.load(e)
if not payload.get("ok"):
err = payload.get("error", {})
raise RuntimeError(f"{err.get('code')}: {err.get('message')}")
return payload["data"]
import { readFileSync } from "node:fs";
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "stats-desk";
// Paste the token from https://stats-desk.skillsafe.ai/tokens.html into a file named "token",
// or replace the fallback with it.
let TOKEN = "YOUR_TOKEN";
try { TOKEN = readFileSync("token", "utf8").trim(); } catch {}
async function call(path, body) {
const res = await fetch(`${BASE}/${path}`, {
method: body ? "POST" : "GET",
headers: {
Authorization: `Bearer ${TOKEN}`,
...(body ? { "Content-Type": "application/json" } : {}),
},
body: body ? JSON.stringify(body) : undefined,
});
const payload = await res.json();
if (!payload.ok) throw new Error(`${payload.error.code}: ${payload.error.message}`);
return payload.data;
}
package main
import (
"bufio"
"bytes"
"crypto/sha256"
"encoding/json"
"fmt"
"io"
"net/http"
"os"
"strings"
"time"
)
const (
base = "https://api.skillsafe.ai/v1/app-api"
slug = "stats-desk"
)
var token = os.Getenv("SKILLSAFE_TOKEN") // from https://stats-desk.skillsafe.ai/tokens.html
type envelope struct {
OK bool `json:"ok"`
Data json.RawMessage `json:"data"`
Error struct {
Code string `json:"code"`
Message string `json:"message"`
} `json:"error"`
}
func call(path string, body any) (json.RawMessage, error) {
method := http.MethodGet
var rdr io.Reader
if body != nil {
method = http.MethodPost
b, _ := json.Marshal(body)
rdr = bytes.NewReader(b)
}
req, _ := http.NewRequest(method, base+"/"+path, rdr)
req.Header.Set("Authorization", "Bearer "+token)
if body != nil {
req.Header.Set("Content-Type", "application/json")
}
res, err := http.DefaultClient.Do(req)
if err != nil {
return nil, err
}
defer res.Body.Close()
var env envelope
if err := json.NewDecoder(res.Body).Decode(&env); err != nil {
return nil, err
}
if !env.OK {
return nil, fmt.Errorf("%s: %s", env.Error.Code, env.Error.Message)
}
return env.Data, nil
}
import java.net.URI;
import java.net.http.*;
public class StatsDesk {
static final String BASE = "https://api.skillsafe.ai/v1/app-api";
static final String SLUG = "stats-desk";
static final String TOKEN = System.getenv().getOrDefault("SKILLSAFE_TOKEN", "YOUR_TOKEN");
static final HttpClient HTTP = HttpClient.newHttpClient();
static String call(String path, String jsonBody) throws Exception {
HttpRequest.Builder b = HttpRequest.newBuilder(URI.create(BASE + "/" + path))
.header("Authorization", "Bearer " + TOKEN);
if (jsonBody != null) {
b.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString(jsonBody));
} else {
b.GET();
}
HttpResponse<String> res = HTTP.send(b.build(), HttpResponse.BodyHandlers.ofString());
// The envelope is always {"ok":true,"data":...} or {"ok":false,"error":...}.
return res.body();
}
}
require "json"
require "net/http"
require "uri"
BASE = "https://api.skillsafe.ai/v1/app-api"
SLUG = "stats-desk"
TOKEN = ENV.fetch("SKILLSAFE_TOKEN", "YOUR_TOKEN") # from https://stats-desk.skillsafe.ai/tokens.html
def call(path, body = nil)
uri = URI("#{BASE}/#{path}")
req = body ? Net::HTTP::Post.new(uri) : Net::HTTP::Get.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
if body
req["Content-Type"] = "application/json"
req.body = JSON.generate(body)
end
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
payload = JSON.parse(res.body)
raise "#{payload['error']['code']}: #{payload['error']['message']}" unless payload["ok"]
payload["data"]
end
<?php
const BASE = "https://api.skillsafe.ai/v1/app-api";
const SLUG = "stats-desk";
define("TOKEN", getenv("SKILLSAFE_TOKEN") ?: "YOUR_TOKEN"); // from /tokens.html
function call(string $path, ?array $body = null) {
$ch = curl_init(BASE . "/" . $path);
$headers = ["Authorization: Bearer " . TOKEN];
if ($body !== null) {
$headers[] = "Content-Type: application/json";
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode($body));
}
curl_setopt($ch, CURLOPT_HTTPHEADER, $headers);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$payload = json_decode(curl_exec($ch), true);
curl_close($ch);
if (empty($payload["ok"])) {
throw new RuntimeException($payload["error"]["code"] . ": " . $payload["error"]["message"]);
}
return $payload["data"];
}
using System.Net.Http.Json;
using System.Text.Json;
static class StatsDesk
{
const string Base = "https://api.skillsafe.ai/v1/app-api";
const string Slug = "stats-desk";
static readonly string Token =
Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN";
static readonly HttpClient Http = new();
public static async Task<JsonElement> Call(string path, object? body = null)
{
var req = new HttpRequestMessage(body is null ? HttpMethod.Get : HttpMethod.Post, $"{Base}/{path}");
req.Headers.Add("Authorization", $"Bearer {Token}");
if (body is not null) req.Content = JsonContent.Create(body);
var res = await Http.SendAsync(req);
var payload = await res.Content.ReadFromJsonAsync<JsonElement>();
if (!payload.GetProperty("ok").GetBoolean())
{
var e = payload.GetProperty("error");
throw new Exception($"{e.GetProperty("code")}: {e.GetProperty("message")}");
}
return payload.GetProperty("data");
}
}
2. Get a token
The easiest route is the token page: it shows the token this browser
already holds, with Copy token and Copy shell export buttons, and
a sign-in button for a personal token. A guest token, minted with
POST /guest and {"slug":"stats-desk"}, can call /me and
/estimate; the run is metered, so /run and /run-stream need
a personal token.
# The token page is the shortest path. It shows the token this browser holds and
# hands you a ready-made shell export:
#
# https://stats-desk.skillsafe.ai/tokens.html
# export SKILLSAFE_TOKEN="..."
#
# To mint a guest token from the command line instead. A guest token is enough
# for /me and /estimate; a run needs a personal token from signing in.
curl -sS -X POST "https://api.skillsafe.ai/v1/app-api/guest" \
-H "Content-Type: application/json" -d '{"slug":"stats-desk"}'
# {"ok":true,"data":{"token":"…","subject_type":"guest"}}
# Open https://stats-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot start a metered run.
import json, urllib.request
req = urllib.request.Request(
"https://api.skillsafe.ai/v1/app-api/guest", data=b'{"slug": "stats-desk"}', method="POST")
req.add_header("Content-Type", "application/json")
with urllib.request.urlopen(req) as r:
TOKEN = json.load(r)["data"]["token"]
// Open https://stats-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
const res = await fetch("https://api.skillsafe.ai/v1/app-api/guest", {
method: "POST",
headers: { "Content-Type": "application/json" },
body: JSON.stringify({ slug: "stats-desk" }),
});
const TOKEN = (await res.json()).data.token;
// Open https://stats-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
guestReq, _ := http.NewRequest(http.MethodPost,
"https://api.skillsafe.ai/v1/app-api/guest", bytes.NewReader([]byte(`{"slug":"stats-desk"}`)))
guestReq.Header.Set("Content-Type", "application/json")
guestRes, err := http.DefaultClient.Do(guestReq)
if err != nil {
panic(err)
}
defer guestRes.Body.Close()
var guest struct {
Data struct {
Token string `json:"token"`
} `json:"data"`
}
_ = json.NewDecoder(guestRes.Body).Decode(&guest)
fmt.Println(guest.Data.Token)
// Open https://stats-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
var http = HttpClient.newHttpClient();
var guestReq = HttpRequest.newBuilder(URI.create("https://api.skillsafe.ai/v1/app-api/guest"))
.header("Content-Type", "application/json")
.POST(HttpRequest.BodyPublishers.ofString("{\"slug\":\"stats-desk\"}"))
.build();
HttpResponse<String> guest = http.send(guestReq, HttpResponse.BodyHandlers.ofString());
System.out.println(guest.body()); // {"ok":true,"data":{"token":"…","subject_type":"guest"}}
# Open https://stats-desk.skillsafe.ai/tokens.html and press "Copy token",
# or mint a guest token here. A guest token can call /me and /estimate but
# cannot start a metered run.
require "json"
require "net/http"
require "uri"
uri = URI("https://api.skillsafe.ai/v1/app-api/guest")
req = Net::HTTP::Post.new(uri)
req["Content-Type"] = "application/json"
req.body = JSON.generate({ slug: "stats-desk" })
res = Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }
TOKEN = JSON.parse(res.body)["data"]["token"]
<?php
// Open https://stats-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
$ch = curl_init("https://api.skillsafe.ai/v1/app-api/guest");
curl_setopt($ch, CURLOPT_POST, true);
curl_setopt($ch, CURLOPT_POSTFIELDS, json_encode(["slug" => "stats-desk"]));
curl_setopt($ch, CURLOPT_HTTPHEADER, ["Content-Type: application/json"]);
curl_setopt($ch, CURLOPT_RETURNTRANSFER, true);
$guest = json_decode(curl_exec($ch), true);
curl_close($ch);
echo $guest["data"]["token"];
// Open https://stats-desk.skillsafe.ai/tokens.html and press "Copy token",
// or mint a guest token here. A guest token can call /me and /estimate but
// cannot start a metered run.
using var http = new HttpClient();
var guestReq = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/guest");
guestReq.Content = new StringContent("{\"slug\":\"stats-desk\"}", Encoding.UTF8, "application/json");
var guestRes = await http.SendAsync(guestReq);
var guest = await guestRes.Content.ReadFromJsonAsync<JsonElement>();
Console.WriteLine(guest.GetProperty("data").GetProperty("token").GetString());
3. Check the session and the balance
call me
# {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}
me = call("me")
print(me["subject_type"], me.get("credits"))
const me = await call("me");
console.log(me.subject_type, me.credits);
raw, err := call("me", nil)
if err != nil {
panic(err)
}
var me struct {
SubjectType string `json:"subject_type"`
Credits int `json:"credits"`
}
_ = json.Unmarshal(raw, &me)
fmt.Println(me.SubjectType, me.Credits)
System.out.println(call("me", null));
// {"ok":true,"data":{"subject_type":"user","username":"you","credits":51234}}
me = call("me")
puts "#{me['subject_type']} #{me['credits']}"
<?php
$me = call("me");
echo $me["subject_type"], " ", $me["credits"], PHP_EOL;
var me = await StatsDesk.Call("me");
Console.WriteLine(me.GetProperty("subject_type").GetString());
4. Price the run (free)
/estimate returns the model binding and the credits a run would reserve. It creates no
job and charges nothing. Expect model_alias gpt-terra and
markup_bps 1000 (a 10% markup). hold_credits is a
reservation, not the price: it is held against your balance while the run executes
and released afterwards. min_credits is the least balance that can start a run. What
you actually pay is charged_credits, reported on the finished job and in the
done event, and it is usually far lower than the hold. The body is the input object
itself, with no {"input": …} wrapper. /estimate does not validate the
body, so check the shape yourself: an object whose every value is a string, task equal
to audit or draft, text and
facts non-empty, and facts a JSON string that parses to an object (this is
what the page's own guard, StatKit.mustBeObject, refuses to spend without).
# body.json is the input object itself - no {"input": ...} wrapper. Build it with
# make-body.js above, or by hand. estimate does not validate it, so check the shape first:
python3 -c 'import json;b=json.load(open("body.json"));assert isinstance(b,dict) and b.get("task") in ("audit","draft") and all(isinstance(v,str) for v in b.values()) and all(b.get(k,"").strip() for k in ("text","facts")) and isinstance(json.loads(b["facts"]),dict)'
INPUT=$(cat body.json)
call estimate "$INPUT"
# {"ok":true,"data":{"model":"...","model_alias":"gpt-terra",
# "markup_bps":1000,"hold_credits":...,"min_credits":...,"sponsor_enabled":false,
# "warnings":[]}}
#
# estimate creates no job and charges nothing. hold_credits is RESERVED, not the
# price; charged_credits after the run is the actual cost, usually far lower.
INPUT = json.load(open("body.json")) # built by make-body.js above, or by hand
assert isinstance(INPUT, dict) and INPUT.get("task") in ("audit", "draft")
assert all(isinstance(v, str) for v in INPUT.values())
assert all(INPUT.get(k, "").strip() for k in ("text", "facts"))
assert isinstance(json.loads(INPUT["facts"]), dict) # facts is a JSON STRING
est = call("estimate", INPUT)
print(est["model_alias"], est["markup_bps"], est["hold_credits"], est.get("warnings"))
me = call("me")
if me.get("credits", 0) < est["min_credits"]:
raise SystemExit("top up first: balance is below min_credits")
const INPUT = JSON.parse(readFileSync("body.json", "utf8")); // built by make-body.js above
if (!INPUT || typeof INPUT !== "object" || !["audit", "draft"].includes(INPUT.task)) throw new Error("task must be audit or draft");
for (const [k, v] of Object.entries(INPUT)) if (typeof v !== "string") throw new Error(k + " must be a string");
for (const k of ["text", "facts"]) if (!(INPUT[k] || "").trim()) throw new Error(k + " is required");
JSON.parse(INPUT.facts); // throws unless facts is a JSON string
const est = await call("estimate", INPUT);
console.log(est.model_alias, est.markup_bps, est.hold_credits, est.warnings);
const me = await call("me");
if ((me.credits ?? 0) < est.min_credits) throw new Error("top up first");
raw, _ := os.ReadFile("body.json") // built by make-body.js above
var input map[string]string // every field is a string, facts included
if err := json.Unmarshal(raw, &input); err != nil {
panic("body.json must be an object of strings: " + err.Error())
}
if input["task"] != "audit" && input["task"] != "draft" {
panic("task must be audit or draft")
}
for _, k := range []string{"text", "facts"} {
if strings.TrimSpace(input[k]) == "" {
panic(k + " is required")
}
}
var facts map[string]any
if err := json.Unmarshal([]byte(input["facts"]), &facts); err != nil {
panic("facts must be a JSON string holding an object")
}
est, err := call("estimate", input)
if err != nil {
panic(err)
}
fmt.Println(string(est)) // model_alias gpt-terra, markup_bps 1000, hold_credits, min_credits
String input = Files.readString(Path.of("body.json")); // built by make-body.js above
if (!input.matches("(?s)\\s*\\{.*\"task\"\\s*:\\s*\"(audit|draft)\".*\\}\\s*"))
throw new IllegalStateException("body.json must be an object with task audit or draft");
String lane = input.replaceAll("(?s).*\"task\"\\s*:\\s*\"(audit|draft)\".*", "$1");
for (String k : new String[] {"text", "facts"})
if (!input.contains("\"" + k + "\"")) throw new IllegalStateException(k + " is required");
String est = call("estimate", input);
System.out.println(est); // model_alias gpt-terra, markup_bps 1000, hold_credits, min_credits
INPUT = JSON.parse(File.read("body.json")) # built by make-body.js above
raise "task must be audit or draft" unless %w[audit draft].include?(INPUT["task"])
INPUT.each { |k, v| raise "#{k} must be a string" unless v.is_a?(String) }
%w[text facts].each { |k| raise "#{k} is required" if INPUT[k].to_s.strip.empty? }
raise "facts must hold an object" unless JSON.parse(INPUT["facts"]).is_a?(Hash)
est = call("estimate", INPUT)
puts est["model_alias"], est["markup_bps"], est["hold_credits"]
<?php
$input = json_decode(file_get_contents("body.json"), true); // built by make-body.js above
if (!is_array($input) || !in_array($input["task"] ?? "", ["audit", "draft"], true)) { throw new Exception("task must be audit or draft"); }
foreach ($input as $k => $v) { if (!is_string($v)) { throw new Exception("$k must be a string"); } }
foreach (["text", "facts"] as $k) { if (trim($input[$k] ?? "") === "") { throw new Exception("$k is required"); } }
if (!is_array(json_decode($input["facts"], true))) { throw new Exception("facts must be a JSON string"); }
$est = call("estimate", $input);
echo $est["model_alias"], " ", $est["markup_bps"], " ", $est["hold_credits"], PHP_EOL;
var input = File.ReadAllText("body.json"); // built by make-body.js above
using var doc = JsonDocument.Parse(input);
var root = doc.RootElement;
var lane = root.GetProperty("task").GetString();
if (lane != "audit" && lane != "draft") throw new Exception("task must be audit or draft");
foreach (var p in root.EnumerateObject())
if (p.Value.ValueKind != JsonValueKind.String) throw new Exception($"{p.Name} must be a string");
JsonDocument.Parse(root.GetProperty("facts").GetString()!); // facts is a JSON string
var est = await StatsDesk.Call("estimate", root);
Console.WriteLine(est); // model_alias gpt-terra, markup_bps 1000, hold_credits, min_credits
5. Run it, then poll
POST /run returns a job_id; poll GET /jobs/{id} until it is
terminal. The reply is a string at data.output.output: JSON.parse
it (step 7). Send an Idempotency-Key built from the lane, a hash of the input and the
attempt number, stats-desk:<lane>:<hash>:a<attempt> (for example
stats-desk:audit:lb7cqsc4tqiq:a1), so a retried request returns the same job instead
of billing a second run. Use one key per distinct input: changed text, changed design notes, changed
facts or a changed audit are a new hash, the same text in the other lane is a new key, and replaying
an old key with a different body is a 409. The page uses
StatKit.hashInput(body) for the hash (it covers task, target,
journal, title, design, text, facts,
audit and question;
make-body.js prints the key); any stable digest of the body works from other
languages. Leave retry_note out of the hash and bump the attempt instead.
# Always send an Idempotency-Key derived from the input. A retried request with
# the same key returns the SAME job instead of billing a second run.
LANE=$(printf '%s' "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["task"])') # audit or draft
KEY="stats-desk:$LANE:$(printf '%s' "$INPUT" | shasum -a 256 | cut -c1-16):a1"
JOB=$(curl -sS -X POST "$BASE/run" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-d "$INPUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["job_id"])')
while :; do
OUT=$(call "jobs/$JOB")
STATUS=$(printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["status"])')
[ "$STATUS" = "succeeded" ] && break
[ "$STATUS" = "failed" ] && echo "$OUT" && exit 1
sleep 2
done
# {"ok":true,"data":{"job_id":"job_...","status":"succeeded",
# "output":{"output":"{\"lane\":\"audit\",\"verdict\":\"needs_revision\",\"headline\":\"...\", ...}"},
# "charged_credits":...,"truncated":false}}
printf '%s' "$OUT" | python3 -c 'import sys,json;print(json.load(sys.stdin)["data"]["output"]["output"])' > reply.json
import hashlib, time
digest = hashlib.sha256(json.dumps(INPUT, sort_keys=True).encode()).hexdigest()[:16]
key = f"stats-desk:{INPUT['task']}:{digest}:a1"
req = urllib.request.Request(f"{BASE}/run", data=json.dumps(INPUT).encode(), method="POST")
req.add_header("Authorization", f"Bearer {TOKEN}")
req.add_header("Content-Type", "application/json")
req.add_header("Idempotency-Key", key)
with urllib.request.urlopen(req) as r:
job_id = json.load(r)["data"]["job_id"]
while True:
job = call(f"jobs/{job_id}")
if job["status"] in ("succeeded", "failed"):
break
time.sleep(2)
if job["status"] == "failed":
raise RuntimeError(job.get("error"))
text = job["output"]["output"] # the reply, as a string
print("charged", job.get("charged_credits"), "truncated", job.get("truncated"))
import { createHash } from "node:crypto";
const digest = createHash("sha256").update(JSON.stringify(INPUT)).digest("hex").slice(0, 16);
const key = `stats-desk:${INPUT.task}:${digest}:a1`;
const started = await fetch(`${BASE}/run`, {
method: "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json", "Idempotency-Key": key },
body: JSON.stringify(INPUT),
}).then((r) => r.json());
if (!started.ok) throw new Error(`${started.error.code}: ${started.error.message}`);
let job = started.data;
while (job.status !== "succeeded" && job.status !== "failed") {
await new Promise((r) => setTimeout(r, 2000));
job = await call(`jobs/${job.job_id}`);
}
if (job.status === "failed") throw new Error(JSON.stringify(job.error));
const text = job.output.output; // the reply, as a string
console.log(job.charged_credits, job.truncated);
body, _ := json.Marshal(input)
sum := sha256.Sum256(body)
key := fmt.Sprintf("stats-desk:%s:%x:a1", input["task"], sum[:8])
req, _ := http.NewRequest(http.MethodPost, base+"/run", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
res, err := http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
var started struct {
Data struct {
JobID string `json:"job_id"`
} `json:"data"`
}
_ = json.NewDecoder(res.Body).Decode(&started)
res.Body.Close()
var jobOutput string
for {
raw, err := call("jobs/"+started.Data.JobID, nil)
if err != nil {
panic(err)
}
var job struct {
Status string `json:"status"`
Output struct {
Output string `json:"output"`
} `json:"output"`
Charged int `json:"charged_credits"`
Truncated bool `json:"truncated"`
}
_ = json.Unmarshal(raw, &job)
if job.Status == "succeeded" {
jobOutput = job.Output.Output
fmt.Println(job.Charged, job.Truncated)
break
}
if job.Status == "failed" {
panic(string(raw))
}
time.Sleep(2 * time.Second)
}
String key = "stats-desk:" + lane + ":" + sha256Hex(input).substring(0, 16) + ":a1";
HttpRequest run = HttpRequest.newBuilder(URI.create(BASE + "/run"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.POST(HttpRequest.BodyPublishers.ofString(input)).build();
String started = HTTP.send(run, HttpResponse.BodyHandlers.ofString()).body();
String jobId = started.replaceAll(".*\"job_id\":\"([^\"]+)\".*", "$1");
while (true) {
String job = call("jobs/" + jobId, null);
if (job.contains("\"status\":\"succeeded\"")) { System.out.println(job); break; }
if (job.contains("\"status\":\"failed\"")) throw new RuntimeException(job);
Thread.sleep(2000);
}
// Parse data.output.output (a string holding the reply JSON) with your JSON library.
// sha256Hex: HexFormat.of().formatHex(MessageDigest.getInstance("SHA-256").digest(input.getBytes(UTF_8)))
require "digest"
key = "stats-desk:#{INPUT['task']}:#{Digest::SHA256.hexdigest(JSON.generate(INPUT))[0, 16]}:a1"
uri = URI("#{BASE}/run")
req = Net::HTTP::Post.new(uri)
req["Authorization"] = "Bearer #{TOKEN}"
req["Content-Type"] = "application/json"
req["Idempotency-Key"] = key
req.body = JSON.generate(INPUT)
job = JSON.parse(Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) { |h| h.request(req) }.body)["data"]
until %w[succeeded failed].include?(job["status"])
sleep 2
job = call("jobs/#{job['job_id']}")
end
raise job.inspect if job["status"] == "failed"
text = job["output"]["output"] # the reply, as a string
puts job["charged_credits"], job["truncated"]
<?php
$key = "stats-desk:" . $input["task"] . ":" . substr(hash("sha256", json_encode($input)), 0, 16) . ":a1";
$ch = curl_init(BASE . "/run");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => json_encode($input),
CURLOPT_HTTPHEADER => ["Authorization: Bearer " . TOKEN, "Content-Type: application/json", "Idempotency-Key: " . $key],
CURLOPT_RETURNTRANSFER => true,
]);
$job = json_decode(curl_exec($ch), true)["data"];
curl_close($ch);
while (!in_array($job["status"], ["succeeded", "failed"], true)) {
sleep(2);
$job = call("jobs/" . $job["job_id"]);
}
if ($job["status"] === "failed") { throw new RuntimeException(json_encode($job)); }
$text = $job["output"]["output"]; // the reply, as a string
echo $job["charged_credits"], PHP_EOL;
using System.Security.Cryptography;
var json = input; // the body.json text from step 4
var key = $"stats-desk:{lane}:" + Convert.ToHexString(SHA256.HashData(System.Text.Encoding.UTF8.GetBytes(json)))[..16].ToLower() + ":a1";
var req = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run");
req.Headers.Add("Authorization", $"Bearer {Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN"}");
req.Headers.Add("Idempotency-Key", key);
req.Content = new StringContent(json, System.Text.Encoding.UTF8, "application/json");
var started = await (await new HttpClient().SendAsync(req)).Content.ReadFromJsonAsync<JsonElement>();
var jobId = started.GetProperty("data").GetProperty("job_id").GetString();
JsonElement job;
while (true)
{
job = await StatsDesk.Call($"jobs/{jobId}");
var status = job.GetProperty("status").GetString();
if (status == "succeeded") break;
if (status == "failed") throw new Exception(job.ToString());
await Task.Delay(2000);
}
var output = job.GetProperty("output").GetProperty("output").GetString()!; // the reply, as a string
6. Or stream it
POST /run-stream takes the same body and headers and answers with server-sent events:
job (the job id), delta (chunks of the reply) and done (the
status, charged_credits, truncated and, when present, the full
output). A browser page may receive only tick heartbeats and then
done, never a delta, so take the reply from done.output.output
when it is there, fall back to the concatenated deltas, and fall back again to
GET /jobs/{id}.
# Server-sent events. `delta` events carry chunks of the reply; `done` carries the
# status, charged_credits and the truncated flag. Ignore `tick` heartbeats.
curl -N -X POST "$BASE/run-stream" \
-H "Authorization: Bearer $TOKEN" \
-H "Content-Type: application/json" \
-H "Idempotency-Key: $KEY" \
-H "Accept: text/event-stream" \
-d "$INPUT"
# event: job {"job_id":"job_..."}
# event: delta {"text":"{\"lane\":\"audit\",\"verdict\":\"needs_revision\",\"headline\":\"The rep"}
# event: done {"status":"succeeded","charged_credits":...,"truncated":false}
req = urllib.request.Request(f"{BASE}/run-stream", data=json.dumps(INPUT).encode(), method="POST")
for h, v in (("Authorization", f"Bearer {TOKEN}"), ("Content-Type", "application/json"),
("Idempotency-Key", key), ("Accept", "text/event-stream")):
req.add_header(h, v)
raw, done, event = "", {}, None
with urllib.request.urlopen(req) as stream:
for line in stream:
line = line.decode().rstrip("\n")
if line.startswith("event: "):
event = line[7:]
elif line.startswith("data: ") and event == "delta":
raw += json.loads(line[6:]).get("text", "")
elif line.startswith("data: ") and event == "done":
done = json.loads(line[6:])
text = (done.get("output") or {}).get("output") or raw
print(done.get("status"), done.get("charged_credits"), done.get("truncated"))
const res = await fetch(`${BASE}/run-stream`, {
method: "POST",
headers: { Authorization: `Bearer ${TOKEN}`, "Content-Type": "application/json", "Idempotency-Key": key, Accept: "text/event-stream" },
body: JSON.stringify(INPUT),
});
const reader = res.body.getReader();
const dec = new TextDecoder();
let buf = "", raw = "", event = null, done = null;
for (;;) {
const { value, done: end } = await reader.read();
if (end) break;
buf += dec.decode(value, { stream: true });
let i;
while ((i = buf.indexOf("\n")) >= 0) {
const line = buf.slice(0, i); buf = buf.slice(i + 1);
if (line.startsWith("event: ")) event = line.slice(7);
else if (line.startsWith("data: ") && event === "delta") raw += JSON.parse(line.slice(6)).text || "";
else if (line.startsWith("data: ") && event === "done") done = JSON.parse(line.slice(6));
}
}
const streamed = done?.output?.output || raw; // browsers may get only ticks + done
console.log(done, streamed.length);
req, _ = http.NewRequest(http.MethodPost, base+"/run-stream", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer "+token)
req.Header.Set("Content-Type", "application/json")
req.Header.Set("Idempotency-Key", key)
req.Header.Set("Accept", "text/event-stream")
res, err = http.DefaultClient.Do(req)
if err != nil {
panic(err)
}
defer res.Body.Close()
var raw strings.Builder
event := ""
sc := bufio.NewScanner(res.Body)
sc.Buffer(make([]byte, 1<<20), 1<<20)
for sc.Scan() {
line := sc.Text()
switch {
case strings.HasPrefix(line, "event: "):
event = line[7:]
case strings.HasPrefix(line, "data: ") && event == "delta":
var d struct{ Text string `json:"text"` }
_ = json.Unmarshal([]byte(line[6:]), &d)
raw.WriteString(d.Text)
case strings.HasPrefix(line, "data: ") && event == "done":
fmt.Println("done:", line[6:])
}
}
HttpRequest stream = HttpRequest.newBuilder(URI.create(BASE + "/run-stream"))
.header("Authorization", "Bearer " + TOKEN)
.header("Content-Type", "application/json")
.header("Idempotency-Key", key)
.header("Accept", "text/event-stream")
.POST(HttpRequest.BodyPublishers.ofString(input)).build();
HTTP.send(stream, HttpResponse.BodyHandlers.ofLines()).body().forEach(line -> {
// "event: delta" lines are followed by "data: {\"text\":...}"; "event: done" by the status.
if (line.startsWith("data: ")) System.out.println(line.substring(6));
});
uri = URI("#{BASE}/run-stream")
req = Net::HTTP::Post.new(uri)
{ "Authorization" => "Bearer #{TOKEN}", "Content-Type" => "application/json",
"Idempotency-Key" => key, "Accept" => "text/event-stream" }.each { |k, v| req[k] = v }
req.body = JSON.generate(INPUT)
raw, event = +"", nil
Net::HTTP.start(uri.hostname, uri.port, use_ssl: true) do |h|
h.request(req) do |res|
res.read_body do |chunk|
chunk.each_line do |line|
line = line.chomp
if line.start_with?("event: ") then event = line[7..]
elsif line.start_with?("data: ") && event == "delta" then raw << JSON.parse(line[6..])["text"].to_s
elsif line.start_with?("data: ") && event == "done" then puts line[6..]
end
end
end
end
end
<?php
$raw = ""; $event = null;
$ch = curl_init(BASE . "/run-stream");
curl_setopt_array($ch, [
CURLOPT_POST => true,
CURLOPT_POSTFIELDS => json_encode($input),
CURLOPT_HTTPHEADER => ["Authorization: Bearer " . TOKEN, "Content-Type: application/json", "Idempotency-Key: " . $key, "Accept: text/event-stream"],
CURLOPT_WRITEFUNCTION => function ($ch, $chunk) use (&$raw, &$event) {
foreach (explode("\n", $chunk) as $line) {
if (str_starts_with($line, "event: ")) $event = substr($line, 7);
elseif (str_starts_with($line, "data: ") && $event === "delta") $raw .= json_decode(substr($line, 6), true)["text"] ?? "";
elseif (str_starts_with($line, "data: ") && $event === "done") echo substr($line, 6), PHP_EOL;
}
return strlen($chunk);
},
]);
curl_exec($ch);
curl_close($ch);
var sreq = new HttpRequestMessage(HttpMethod.Post, "https://api.skillsafe.ai/v1/app-api/run-stream");
sreq.Headers.Add("Authorization", $"Bearer {Environment.GetEnvironmentVariable("SKILLSAFE_TOKEN") ?? "YOUR_TOKEN"}");
sreq.Headers.Add("Idempotency-Key", key);
sreq.Headers.Add("Accept", "text/event-stream");
sreq.Content = new StringContent(json, System.Text.Encoding.UTF8, "application/json");
using var sres = await new HttpClient().SendAsync(sreq, HttpCompletionOption.ResponseHeadersRead);
using var sr = new StreamReader(await sres.Content.ReadAsStreamAsync());
var raw = new System.Text.StringBuilder(); string? ev = null, line;
while ((line = await sr.ReadLineAsync()) != null)
{
if (line.StartsWith("event: ")) ev = line[7..];
else if (line.StartsWith("data: ") && ev == "delta") raw.Append(JsonSerializer.Deserialize<JsonElement>(line[6..]).GetProperty("text").GetString());
else if (line.StartsWith("data: ") && ev == "done") Console.WriteLine(line[6..]);
}
7. Parse the reply
The reply is one JSON object, delivered as a string in data.output.output; you must
JSON.parse it. The model is told to send no code fences, but tolerate them: strip a
leading ```json and a trailing ```, keep everything from the first
{ to the last }, and parse that outer object. Then branch on
lane: the envelope keys are the same for both lanes, the body keys are not.
# reply.json holds data.output.output from step 5. Strip any fence, keep the object:
python3 - <<'EOF'
import json, re
t = open("reply.json").read().strip()
t = re.sub(r"^```(?:json)?\s*", "", t, flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
r = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(r["lane"], r["verdict"], "-", r["headline"])
for i in r["issues"]:
print(i["id"], i["priority"], i["ref"] or "-", "L%s" % i["line"], i["category"], "|", i["problem"])
if r["lane"] == "audit":
print("UNIT", r["scope"]["unit"])
for c in r["claims"]:
print("CLAIM L%s" % c["line"], c["status"], c["claim"], "|", c["analysis"])
open("revision.txt", "w").write(r["revision"]["text"])
else:
open("statistical-analysis.txt", "w").write(r["draft_text"])
print("UNRESOLVED", len(r["reporting_notes"]["unresolved"]))
print("LEGEND", r["legend_template"])
for q in r["author_input_needed"]:
print("ASK", q)
EOF
import re
t = re.sub(r"^```(?:json)?\s*", "", text.strip(), flags=re.I)
t = re.sub(r"\s*```\s*$", "", t)
reply = json.loads(t[t.index("{"):t.rindex("}") + 1])
print(reply["lane"], reply["verdict"], reply["headline"])
print([(i["id"], i["priority"], i["ref"], i["line"]) for i in reply["issues"]])
if reply["lane"] == "audit":
print(reply["scope"]["unit"])
print([(c["line"], c["status"], c["claim"]) for c in reply["claims"]])
print(reply["revision"]["section"], reply["revision"]["text"][:200])
else:
print(reply["draft_text"][:200])
print(reply["reporting_notes"]["unresolved"])
print(reply["author_input_needed"])
let t = text.trim().replace(/^```(?:json)?\s*/i, "").replace(/\s*```\s*$/, "");
const reply = JSON.parse(t.slice(t.indexOf("{"), t.lastIndexOf("}") + 1));
console.log(reply.lane, reply.verdict, reply.headline);
console.log(reply.issues.map((i) => `${i.id} ${i.priority} [${i.ref}] L${i.line}: ${i.problem}`));
if (reply.lane === "audit") console.log(reply.scope.unit, reply.claims.map((c) => [c.line, c.status, c.claim]), reply.revision.section);
else console.log(reply.draft_text.length, "chars;", reply.reporting_notes.unresolved.length, "unresolved;", reply.legend_template);
console.log(reply.author_input_needed);
t := strings.TrimSpace(jobOutput)
t = strings.TrimPrefix(strings.TrimPrefix(t, "```json"), "```")
t = strings.TrimSuffix(strings.TrimSpace(t), "```")
start, end := strings.Index(t, "{"), strings.LastIndex(t, "}")
var reply struct {
Lane string `json:"lane"`
Verdict string `json:"verdict"`
Headline string `json:"headline"`
TLDR []string `json:"tldr"`
Issues []struct {
ID, Ref, Priority, Category string
Line int
Evidence, Problem, Why, Fix string
} `json:"issues"`
// audit lane
Scope struct {
InputReviewed string `json:"input_reviewed"`
Boundary string `json:"boundary"`
Design string `json:"design"`
Unit string `json:"unit"`
} `json:"scope"`
Claims []struct {
Line int
Claim, Analysis, Status, Note string
NUnit string `json:"n_unit"`
} `json:"claims"`
Revision struct {
Section string `json:"section"`
Text string `json:"text"`
} `json:"revision"`
ReviewerRisk []string `json:"reviewer_risk"`
// draft lane
DraftText string `json:"draft_text"`
ReportingNotes struct {
NDefinition string `json:"n_definition"`
TestsModels string `json:"tests_models"`
MultipleComparisons string `json:"multiple_comparisons"`
Software string `json:"software"`
Tails string `json:"tails"`
Unresolved []string `json:"unresolved"`
} `json:"reporting_notes"`
LegendTemplate string `json:"legend_template"`
// both
AuthorInputNeeded []string `json:"author_input_needed"`
PrescanResponses []struct {
Ref, Verdict, Note string
} `json:"prescan_responses"`
}
if err := json.Unmarshal([]byte(t[start:end+1]), &reply); err != nil {
panic(err)
}
fmt.Println(reply.Lane, reply.Verdict, reply.Headline)
for _, i := range reply.Issues {
fmt.Println(i.ID, i.Priority, i.Ref, i.Line, i.Problem)
}
switch reply.Lane {
case "audit":
fmt.Println(len(reply.Claims), "claims; unit:", reply.Scope.Unit)
case "draft":
fmt.Println(len(reply.ReportingNotes.Unresolved), "placeholders unresolved")
}
// output is data.output.output from step 5: a string holding the reply JSON.
String t = output.strip().replaceFirst("^```(?:json)?\\s*", "").replaceFirst("\\s*```\\s*$", "");
String json = t.substring(t.indexOf('{'), t.lastIndexOf('}') + 1);
// With Jackson: JsonNode r = new ObjectMapper().readTree(json);
// r.get("lane"): audit | draft; r.get("verdict"): not_ready | needs_revision | ready
// r.get("issues"), r.get("author_input_needed"), r.get("prescan_responses")
// audit: r.get("scope"), r.get("claims"), r.get("revision"), r.get("reviewer_risk")
// draft: r.get("draft_text"), r.get("reporting_notes"), r.get("legend_template")
System.out.println(json);
t = text.strip.sub(/\A```(?:json)?\s*/i, "").sub(/\s*```\s*\z/, "")
reply = JSON.parse(t[t.index("{")..t.rindex("}")])
puts reply["lane"], reply["verdict"], reply["headline"]
reply["issues"].each { |i| puts "#{i['id']} #{i['priority']} [#{i['ref']}] L#{i['line']}: #{i['problem']}" }
case reply["lane"]
when "audit"
puts "unit: #{reply['scope']['unit']}"
reply["claims"].each { |c| puts "L#{c['line']} #{c['status']} #{c['claim']}" }
File.write("revision.txt", reply["revision"]["text"])
else
File.write("statistical-analysis.txt", reply["draft_text"])
puts "#{reply['reporting_notes']['unresolved'].length} unresolved"
end
reply["author_input_needed"].each { |q| puts "ASK #{q}" }
<?php
$t = preg_replace('/\s*```\s*$/', "", preg_replace('/^```(?:json)?\s*/i', "", trim($text)));
$reply = json_decode(substr($t, strpos($t, "{"), strrpos($t, "}") - strpos($t, "{") + 1), true);
echo $reply["lane"], " ", $reply["verdict"], " ", $reply["headline"], PHP_EOL;
foreach ($reply["issues"] as $i) { echo $i["id"], " ", $i["priority"], " [", $i["ref"], "] L", $i["line"], ": ", $i["problem"], PHP_EOL; }
if ($reply["lane"] === "audit") {
echo "unit: ", $reply["scope"]["unit"], PHP_EOL;
foreach ($reply["claims"] as $c) { echo "L", $c["line"], " ", $c["status"], " ", $c["claim"], PHP_EOL; }
file_put_contents("revision.txt", $reply["revision"]["text"]);
} else {
file_put_contents("statistical-analysis.txt", $reply["draft_text"]);
echo count($reply["reporting_notes"]["unresolved"]), " unresolved", PHP_EOL;
}
foreach ($reply["author_input_needed"] as $q) { echo "ASK ", $q, PHP_EOL; }
var t = System.Text.RegularExpressions.Regex.Replace(output.Trim(), @"^```(?:json)?\s*", "");
t = System.Text.RegularExpressions.Regex.Replace(t, @"\s*```\s*$", "");
var json2 = t.Substring(t.IndexOf('{'), t.LastIndexOf('}') - t.IndexOf('{') + 1);
using var parsed = JsonDocument.Parse(json2);
var r = parsed.RootElement;
Console.WriteLine($"{r.GetProperty("lane")} {r.GetProperty("verdict")} {r.GetProperty("headline")}");
foreach (var i in r.GetProperty("issues").EnumerateArray())
Console.WriteLine($"{i.GetProperty("id")} {i.GetProperty("priority")} [{i.GetProperty("ref")}] L{i.GetProperty("line")}: {i.GetProperty("problem")}");
switch (r.GetProperty("lane").GetString())
{
case "audit": File.WriteAllText("revision.txt", r.GetProperty("revision").GetProperty("text").GetString()); break;
case "draft": File.WriteAllText("statistical-analysis.txt", r.GetProperty("draft_text").GetString()); break;
}
Invariants worth asserting
The web page holds every reply to the browser's facts and to your text before it shows it
(recon.js: normalize(), then reconcile()). Do the same before
you hold a manuscript on a verdict or paste the drafted text:
- It parses: the reply is one JSON object (after stripping any fence), and every key of the lane's contract is present: 11 keys for
audit, 10 fordraft. - Lane:
laneequals thetaskyou sent. A different lane means the model answered the other contract. - Every flag answered once: every flag in
facts.flagshas aprescan_responsesentry (arefmay list several ids, comma-separated); no response names a flag that was not sent; everyverdictisconfirmedordismissed. Read each dismissal and decide whether you agree: the page lists them as items to review. - Confirmed high/medium flags are carried (audit): every
highormediumflag left standing (not dismissed) appears in therefof an issue, and no issue'srefnames a flag the prescan never raised (both lanes). - Verdict floor (audit): one of
not_ready,needs_revision,ready; never looser than the flags left standing (anyhighmeansnot_ready, else anymediummeans at leastneeds_revision), and equal to what the worst issue implies:P0meansnot_ready,P1meansneeds_revision, otherwiseready. - Verdict (draft):
readyonly with noAUTHOR_INPUT_NEEDEDindraft_textorlegend_templateand no open issue; any openP0issue meansnot_ready. - Quotes are on the cited lines: every non-empty
evidenceappears, whitespace and case ignored, on itsline(or the line next to it) of your original text, and nolineinissuesorclaimsis past the end of the text.line: 0withevidence: ""is how the model reports an absence. - No invented numbers: every number that carries a claim in
revision.text,draft_textandlegend_template(every decimal, every integer aftern =, every df insidet(..),F(..),r(..),chi2(..)) is in your text or design notes. The hint in brackets after a placeholder,AUTHOR_INPUT_NEEDED (averaged per mouse, or ...), is a question and is not checked. - No invented methods: every test, correction and software named in that ready-to-paste text is named in your text or design notes (a reported
t(df)orchi2(df)counts as naming the t-test or chi-square test). A method the audit recommends is not one the authors used. - No stray p values: every p value in the prose (headline, TL;DR, issues, claims, flag notes) is one reported in your text, a conventional threshold (0.05, 0.01, 0.001, 0.1, 0.005, 0.0001), or within about 6% of a
p_computedorp_rangebound fromfacts.stats; and everyTn it cites exists infacts.stats. - Placeholders are backed: if the ready-to-paste text carries any
AUTHOR_INPUT_NEEDED,author_input_neededis not empty.
The page's checks are plain JavaScript, so the exact same reconciliation runs in Node:
// check-reply.js - hold a reply to the same checks the web page runs (recon.js reconcile()).
// Save recon.js and statkit.js next to this file.
// Usage: node check-reply.js body.json reply.json (reply.json = data.output.output)
const fs = require("fs");
const K = require("./statkit.js");
const R = require("./recon.js");
const body = JSON.parse(fs.readFileSync(process.argv[2] || "body.json", "utf8"));
const facts = JSON.parse(body.facts);
// Rebuild the prescan from the numbered text the model saw: quotes are checked against these lines.
// (With clipped text, run this on the original text instead so line numbers stay exact.)
const text = body.text.split("\n").map((l) => l.replace(/^\d+\| /, "")).join("\n");
const A = K.analyze({ lane: body.task, target: body.target, journal: body.journal, title: body.title, design: body.design, text: text });
if (JSON.stringify(A.flags.map((f) => f.id)) !== JSON.stringify(facts.flags.map((f) => f.id)))
console.warn("note: facts.flags differ from a fresh prescan of this text; the checks use the fresh one");
const reply = R.normalize(R.parseResult(fs.readFileSync(process.argv[3] || "reply.json", "utf8")), body.task);
const missing = R.SECTIONS[reply.lane].filter((k) => reply.present.indexOf(k) === -1);
if (missing.length) console.log("BAD keys missing: " + missing.join(", "));
const rec = R.reconcile(reply, { analysis: A, input: body });
for (const i of rec.items) console.log((i.bad ? "BAD " : i.review ? "LOOK " : "ok ") + i.kind + ": " + i.text);
const look = rec.items.filter((i) => i.review).length; // dismissals and placeholders: review by hand
console.log(reply.lane, reply.verdict, "-", rec.disagreements + missing.length, "disagreement(s),", look, "to review");
process.exit(rec.disagreements + missing.length ? 1 : 0);
Run on the worked examples below, both the audit and the draft replies report 0 disagreements, with one item to review each: the placeholders left for the authors to fill.
The output contract
Every key of the lane's contract is always present. Arrays may be empty ([], never a
filler such as "None"). An enum is written "a|b|c": the reply carries
exactly one of the values. Text is plain: no Markdown, no emoji, each string under 500 characters
except revision.text and draft_text (under 3,500). The reply is only the
JSON object, with no prose before or after it.
The common envelope (both lanes)
{"lane":"audit|draft",
"verdict":"not_ready|needs_revision|ready",
"headline":"...",
"tldr":["...","..."],
"issues":[{"id":"I1","ref":"S2","priority":"P0|P1|P2","category":"...","line":12,"evidence":"...","problem":"...","why":"...","fix":"..."}],
...the lane body...,
"author_input_needed":["..."],
"prescan_responses":[{"ref":"S1","verdict":"confirmed|dismissed","note":"..."}]}
| key | shape | what it holds |
|---|---|---|
lane | string | The lane answered: your task, or the closest lane when task was missing or unknown. |
verdict | enum | See verdict and priority. In the draft lane it describes the drafted section. |
headline | string | One sentence on the state of the statistics. |
tldr | array of strings | 2-5 bullets. When you sent a question, one bullet starts "Answer:". |
issues | array of objects | Ids I1, I2, ... worst first. ref: the flag ids the issue answers, comma-separated, or "" for one the browser missed. category: snake_case, the flag's category when it answers one. line and evidence: one input line and a verbatim part of it (without the 12| prefix), or 0 and "" for an absence. why: what a statistical reviewer would challenge. Any p value discussed comes from facts.stats, cited by id. In the draft lane: only what the draft cannot resolve on its own. |
author_input_needed | array of strings | Short factual questions, one per missing fact, each naming where the answer goes ("Which software and version ran the mixed model? (Statistical analysis)"). Every AUTHOR_INPUT_NEEDED placeholder is backed by one. |
prescan_responses | array of {ref, verdict, note} | Exactly one per flag in facts.flags, in both lanes, with the reason and the line. Empty when no flags were sent. |
The audit body
{"scope":{"input_reviewed":"...","boundary":"...","design":"...","unit":"..."},
"claims":[{"line":6,"claim":"...","analysis":"...","n_unit":"...","status":"supported|underreported|overclaimed|not_assessable","note":"..."}],
"revision":{"section":"statistical_analysis|results|legend|mixed","text":"...ready-to-paste text with AUTHOR_INPUT_NEEDED..."},
"reviewer_risk":["..."]}
| key | what it holds |
|---|---|
scope | input_reviewed: what was supplied. boundary: what could not be assessed and why. design: the study design as the text states it (groups, treatments, time points, endpoints, paired or repeated structure). unit: the independent experimental unit and how biological replicates, technical replicates and subsamples relate to n, or "not stated". |
claims | At most 12, one per central result claim: the claim in a few words, the test or model reported for it ("not stated" if none), the n and its unit as reported, and a status: supported (analysis, n and uncertainty reported and matching the claim), underreported (may hold, key reporting missing), overclaimed (wording beyond the evidence) or not_assessable. |
revision | Ready-to-paste revised text for the part that most needs it. It keeps every supplied fact, removes overclaiming, and writes AUTHOR_INPUT_NEEDED for every missing fact; no number, test, correction or software appears that is not already in the text or design notes. |
reviewer_risk | 1-4 sentences on what a statistical reviewer may still challenge after the revision. |
The draft body
{"draft_text":"Statistical analysis. ...one string, AUTHOR_INPUT_NEEDED wherever a fact is missing...",
"reporting_notes":{"n_definition":"...","tests_models":"...","multiple_comparisons":"...","software":"...","tails":"...","unresolved":["..."]},
"legend_template":"...one sentence of figure-legend statistics text, with placeholders..."}
| key | what it holds |
|---|---|
draft_text | The subsection as one string, conservative, in the order software, summary convention, sample size and replication, tests or models, multiplicity, exact reporting and thresholds, exclusions. A correction or model the audit recommends but the authors have not confirmed is an AUTHOR_INPUT_NEEDED, not a statement. |
reporting_notes | One short string each for n_definition, tests_models, multiple_comparisons, software and tails, saying what the draft states and where it came from (line numbers or "design notes"); unresolved lists each placeholder left in the draft. |
legend_template | One sentence of figure-legend statistics (what points and bars show, n and its unit, test, correction, exact p values) in the same conservative style. |
Verdict and priority
| value | meaning |
|---|---|
not_ready | Audit: any issue is P0. Draft: the independent unit, or the test for a central claim, is unknown (any open P0). |
needs_revision | Audit: the worst issue is P1. Draft: placeholders remain only for P1 or P2 facts. |
ready | Audit: nothing worse than P2. Draft: no placeholders and no open issue. |
P0 issue | Must fix before submission: the independent unit wrong or undefined for a central claim, a paired, repeated or nested structure ignored, many comparisons with no correction or family, an interaction inferred from a difference in significance, undisclosed exclusions that could change the result, a reported p inconsistent with its statistic on the other side of 0.05, or an analysis that cannot be understood from the text. |
P1 issue | Important: exact n missing, error bars or summary convention undefined, no effect size or uncertainty for key results, threshold-only p values, software missing, a p inconsistent with its statistic on the same side of 0.05, a small-sample limitation unacknowledged, trend language. |
P2 issue | Clarity: inconsistent replicate terminology, panel-specific n missing, unclear ns labels, scattered statistical text. |
The model may be stricter than facts.browser_verdict, never looser, unless it dismissed
the flags that set it. The priority of an issue may also be stricter than its flag's
(significance_difference is a medium flag, but an interaction inferred from a difference
in significance is a P0).
Enums
| where | values | the page's fallback |
|---|---|---|
lane | audit, draft | the lane asked |
verdict | not_ready, needs_revision, ready | derived from the worst issue |
issues[].priority | P0, P1, P2 | P1 |
prescan_responses[].verdict | confirmed, dismissed | confirmed |
claims[].status | supported, underreported, overclaimed, not_assessable | not_assessable |
revision.section | statistical_analysis, results, legend, mixed | mixed |
The page also drops an issue with no problem, fix or evidence,
a claim with no claim, and a flag response with no ref, and reads a
line such as "L12" as 12 (anything unreadable as 0).
Worked example: audit
The page's "Psychology experiment, APA style" example (title Feedback timing and workload,
target general, no design notes): a Statistical analysis paragraph and a Results section
for a three-condition feedback study. The prescan recomputed five test results. Four check out
(T2 to T5: two t-tests, a chi-square and a correlation), but
T1, F(2, 87) = 3.12, p = .012, implies p = 0.0491 (range 0.0489 to 0.0494),
so it is inconsistent, on the same side of 0.05. It raised three flags: S1
stat_inconsistent (medium), S2 trend_language (medium, line 11)
and S3 randomization_blinding (low), so browser_verdict is
needs_revision. There is no multiple_comparisons flag because the text names
Bonferroni. This request is complete and sendable as make-body.js builds it; its
Idempotency-Key from the page is stats-desk:audit:lb7cqsc4tqiq:a1.
The request body, with text and facts abridged:
{
"task": "audit",
"target": "general",
"journal": "",
"title": "Feedback timing and workload",
"design": "",
"text": "1| Statistical analysis\n2| Analyses were conducted in R 4.4.1. All tests were two-tailed with alpha = .05. Continuous ou... [abridged here: all 11 lines are shown decoded below]",
"facts": "{\"target\":\"general\",\"browser_verdict\":\"needs_revision\",\"flags\":[{\"id\":\"S1\",\"severity\":\"medium\",\"priority\":\"P1\",\"category... [abridged here: decoded below]",
"question": ""
}
Its text, decoded (the N| prefixes are part of the string):
1| Statistical analysis
2| Analyses were conducted in R 4.4.1. All tests were two-tailed with alpha = .05. Continuous outcomes are reported as mean ± SD. Pairwise comparisons after the omnibus ANOVA used Bonferroni correction.
3|
4| Results
5| The final sample comprised 90 participants (n = 30 per condition); two participants who failed the attention check were excluded before analysis.
6| Perceived workload differed across the three feedback conditions, F(2, 87) = 3.12, p = .012, partial eta squared = .07.
7| Participants in the delayed-feedback condition reported higher workload than those in the immediate condition (M = 5.1, SD = 1.2 vs M = 4.4, SD = 1.3), t(58) = 2.17, p = .034, d = 0.56.
8| Workload did not differ between the immediate and no-feedback conditions, t(58) = 0.82, p = .42.
9| Condition was associated with task abandonment, chi2(2, N = 90) = 6.35, p = .042.
10| Self-reported effort correlated with workload, r(88) = .31, p = .003.
11| Accuracy was marginally significant in the delayed condition (p = .07), suggesting that delayed feedback drives poorer performance.
Its facts, decoded:
{
"target": "general",
"browser_verdict": "needs_revision",
"flags": [
{
"id": "S1",
"severity": "medium",
"priority": "P1",
"category": "stat_inconsistent",
"message": "T1 (line 6): F(2, 87) = 3.12, p = .012 implies p = 0.0491 (range 0.0489 to 0.0494 allowing for rounding), not p = .012. Check the statistic, the df and the p value.",
"lines": [
6
]
},
{
"id": "S2",
"severity": "medium",
"priority": "P1",
"category": "trend_language",
"message": "Trend language (\"marginally significant\", \"a trend toward significance\") recasts a non-significant result as a weak positive. Report the estimate and interval instead.",
"lines": [
11
]
},
{
"id": "S3",
"severity": "low",
"priority": "P2",
"category": "randomization_blinding",
"message": "Animals or participants are compared, and the text says nothing about randomization or blinding.",
"lines": []
}
],
"stats": [
{
"id": "T1",
"line": 6,
"test": "F",
"df1": 2,
"df2": 87,
"value": 3.12,
"reported": "F(2, 87) = 3.12, p = .012",
"p_reported": "p = .012",
"p_computed": 0.04913,
"p_range": [
0.04891,
0.04936
],
"status": "inconsistent"
},
{
"id": "T2",
"line": 7,
"test": "t",
"df1": 58,
"df2": null,
"value": 2.17,
"reported": "t(58) = 2.17, p = .034",
"p_reported": "p = .034",
"p_computed": 0.03412,
"p_range": [
0.03372,
0.03451
],
"status": "consistent"
},
{
"id": "T3",
"line": 8,
"test": "t",
"df1": 58,
"df2": null,
"value": 0.82,
"reported": "t(58) = 0.82, p = .42",
"p_reported": "p = .42",
"p_computed": 0.4156,
"p_range": [
0.4128,
0.4184
],
"status": "consistent"
},
{
"id": "T4",
"line": 9,
"test": "chi2",
"df1": 2,
"df2": null,
"value": 6.35,
"reported": "chi2(2, N = 90) = 6.35, p = .042",
"p_reported": "p = .042",
"p_computed": 0.04179,
"p_range": [
0.04169,
0.0419
],
"status": "consistent"
},
{
"id": "T5",
"line": 10,
"test": "r",
"df1": 88,
"df2": null,
"value": 0.31,
"reported": "r(88) = .31, p = .003",
"p_reported": "p = .003",
"p_computed": 0.002945,
"p_range": [
0.002495,
0.003466
],
"status": "consistent"
}
],
"n_mentions": [
{
"line": 5,
"text": "n = 30 per condition); two particip",
"unit": "unclear",
"unit_word": "",
"range": false
},
{
"line": 9,
"text": "N = 90) = 6.35, p = .042.",
"unit": "unclear",
"unit_word": "",
"range": false
}
],
"tests_named": [
"ANOVA"
],
"corrections_named": [
"Bonferroni"
],
"software_named": [
"R"
],
"summary_conventions": [
"mean ± s.d."
],
"named_in_design_notes": {
"tests": [],
"corrections": [],
"software": []
},
"tails_stated": true,
"one_tailed_declared": false,
"p_counts": {
"exact": 6,
"threshold": 0,
"ns": 0
},
"clipped": null
}
The reply (data.output.output, parsed; shown in full apart from one phrase in reviewer_risk abbreviated with "..."). The model confirmed all three
flags and carried each in an issue. It added three issues the browser missed (ref: ""),
the most useful being I2: line 2 says the pairwise comparisons were Bonferroni-corrected,
yet p = .034 is the unadjusted value T2 implies, so it is not stated whether
the difference survives the correction. The worst issue is P1, so the verdict is
needs_revision, matching the browser's floor. The revision rewrites line 11 as not
statistically significant, drops "drives", keeps every reported number, and leaves the F value and p
value of line 6 as placeholders rather than guessing which is wrong. check-reply.js
reports 0 disagreements and one item to review: the placeholders.
{
"lane": "audit",
"verdict": "needs_revision",
"headline": "The reporting is mostly complete and the degrees of freedom match the stated sample, but the omnibus F test's p value does not match its statistic, the Bonferroni status of the pairwise p values is unclear, and the accuracy sentence uses trend and causal language without a reported test.",
"tldr": [
"T1 (line 6): F(2, 87) = 3.12 implies p of about 0.0491 (range 0.0489 to 0.0494), not the reported p = .012; the F value, df or p value needs rechecking.",
"Line 2 says pairwise comparisons used Bonferroni correction, but the p = .034 at line 7 matches the unadjusted value implied by the statistic (T2), so it is not stated whether the delayed versus immediate difference survives the stated correction.",
"Line 11 calls p = .07 marginally significant and infers that delayed feedback drives poorer performance, with no test, statistic or n reported for accuracy.",
"The t-tests, the chi-square test and the correlation are not described in the Statistical analysis section, and effect sizes or intervals are missing for lines 8 and 9."
],
"issues": [
{
"id": "I1",
"ref": "S1",
"priority": "P1",
"category": "stat_inconsistent",
"line": 6,
"evidence": "F(2, 87) = 3.12, p = .012, partial eta squared = .07.",
"problem": "The reported p = .012 is inconsistent with F(2, 87) = 3.12; the statistic implies p of about 0.0491 (T1, range 0.0489 to 0.0494 allowing for rounding).",
"why": "A reviewer will ask which of F, df or p is wrong; both values fall below .05, but the implied p lies very close to the threshold, which weakens the omnibus result that licenses the pairwise tests.",
"fix": "Recheck the ANOVA output in R and report the F value, df and exact p value exactly as produced."
},
{
"id": "I2",
"ref": "",
"priority": "P1",
"category": "multiple_comparisons",
"line": 7,
"evidence": "t(58) = 2.17, p = .034, d = 0.56.",
"problem": "Line 2 states that pairwise comparisons used Bonferroni correction, yet p = .034 matches the unadjusted two-tailed p implied by t(58) = 2.17 (T2, p_computed 0.0341), and only two of the three possible pairwise comparisons are reported.",
"why": "A reviewer will ask whether the reported p values are Bonferroni-adjusted, how many comparisons formed the family, and whether the delayed versus immediate difference remains significant after the stated correction; the delayed versus no-feedback comparison is also missing.",
"fix": "State whether the pairwise p values are adjusted or unadjusted, define the family of comparisons, report adjusted p values if the correction was applied, and report the delayed versus no-feedback comparison."
},
{
"id": "I3",
"ref": "S2",
"priority": "P1",
"category": "trend_language",
"line": 11,
"evidence": "Accuracy was marginally significant in the delayed condition (p = .07), suggesting that delayed feedback drives poorer performance.",
"problem": "A non-significant result (p = .07 against the stated alpha = .05) is described as marginally significant and used to infer that delayed feedback drives poorer performance; no test, statistic, comparison group, n or estimate is reported for accuracy.",
"why": "A reviewer will challenge recasting a non-significant result as a weak positive and drawing a causal conclusion from it; the analysis behind the p value cannot be identified from the text.",
"fix": "Report the accuracy test, statistic, df, comparison groups, estimate and exact p value, describe the result as not statistically significant, and remove the causal inference."
},
{
"id": "I4",
"ref": "",
"priority": "P1",
"category": "tests_not_described",
"line": 2,
"evidence": "Pairwise comparisons after the omnibus ANOVA used Bonferroni correction.",
"problem": "The Statistical analysis section names only the omnibus ANOVA; the type of pairwise t-test, the chi-square test at line 9, the correlation coefficient at line 10 and the accuracy analysis at line 11 are not described.",
"why": "A reviewer cannot tell whether the pairwise tests were independent-samples t-tests or post hoc contrasts using the pooled ANOVA error, which correlation coefficient was used, or how accuracy was analysed.",
"fix": "Name each test in the Statistical analysis section, including the pairwise test type, the chi-square test for abandonment, the correlation coefficient type and the accuracy analysis."
},
{
"id": "I5",
"ref": "",
"priority": "P1",
"category": "effect_size_uncertainty",
"line": 8,
"evidence": "Workload did not differ between the immediate and no-feedback conditions, t(58) = 0.82, p = .42.",
"problem": "No effect size or interval is given for the null comparison at line 8 or for the chi-square result at line 9, and no confidence intervals are reported for any key result; line 8 also states no difference from a non-significant test.",
"why": "Without an estimate and interval, a non-significant result cannot be read as evidence of no difference, and the size of the abandonment association cannot be judged.",
"fix": "Report effect sizes with confidence intervals for lines 8 and 9 (and ideally for all key results), and word line 8 as no statistically significant difference."
},
{
"id": "I6",
"ref": "S3",
"priority": "P2",
"category": "randomization_blinding",
"line": 0,
"evidence": "",
"problem": "The text does not state how participants were assigned to the three feedback conditions.",
"why": "A reviewer will ask whether assignment was random, since between-condition differences are interpreted as effects of feedback timing.",
"fix": "State how participants were allocated to conditions, or state that allocation was not randomized."
}
],
"scope": {
"input_reviewed": "A Statistical analysis paragraph (line 2) and a Results section (lines 5 to 11) for a three-condition feedback timing study; no design notes, figures or legends were supplied.",
"boundary": "Raw data, figures and legends were not supplied, so ANOVA assumptions, the source of the T1 mismatch, whether pairwise p values were adjusted, and the accuracy analysis are not assessable from the supplied material.",
"design": "Three feedback conditions (delayed, immediate, no feedback) with 30 participants each after two attention-check exclusions; outcomes are perceived workload, task abandonment, self-reported effort and accuracy. The text suggests a between-participants design (t(58) and F(2, 87) are consistent with separate groups of 30), but the design and assignment method are not stated explicitly.",
"unit": "The participant appears to be the independent unit (n = 30 per condition, N = 90), consistent with the reported degrees of freedom; no technical replicates or subsamples are described."
},
"claims": [
{
"line": 6,
"claim": "Workload differed across the three feedback conditions",
"analysis": "omnibus ANOVA",
"n_unit": "90 participants, 30 per condition",
"status": "underreported",
"note": "The reported p = .012 is inconsistent with F(2, 87) = 3.12 (T1)."
},
{
"line": 7,
"claim": "Delayed feedback produced higher workload than immediate feedback",
"analysis": "t-test, type not stated; Bonferroni correction stated at line 2",
"n_unit": "30 participants per condition",
"status": "underreported",
"note": "p = .034 matches the unadjusted value (T2); whether it survives the stated Bonferroni correction is not stated."
},
{
"line": 8,
"claim": "No workload difference between immediate and no-feedback conditions",
"analysis": "t-test, type not stated",
"n_unit": "30 participants per condition",
"status": "underreported",
"note": "No effect size or interval; a non-significant test is worded as no difference."
},
{
"line": 9,
"claim": "Condition was associated with task abandonment",
"analysis": "chi-square test, not described in the Statistical analysis section",
"n_unit": "N = 90 participants",
"status": "underreported",
"note": "Statistic and p are consistent (T4), but no effect size is reported."
},
{
"line": 10,
"claim": "Self-reported effort correlated with workload",
"analysis": "correlation, coefficient type not stated",
"n_unit": "90 participants implied by r(88)",
"status": "underreported",
"note": "Statistic and p are consistent (T5); the correlation type and interval are not reported."
},
{
"line": 11,
"claim": "Delayed feedback drives poorer accuracy",
"analysis": "not stated",
"n_unit": "not stated",
"status": "overclaimed",
"note": "p = .07 is not significant at the stated alpha, and the causal wording goes beyond the reported evidence."
}
],
"revision": {
"section": "mixed",
"text": "Statistical analysis\nAnalyses were conducted in R 4.4.1. All tests were two-tailed with alpha = .05. Continuous outcomes are reported as mean ± SD. Perceived workload was compared across the three feedback conditions with an omnibus ANOVA (AUTHOR_INPUT_NEEDED: ANOVA design). Pairwise comparisons after the omnibus ANOVA used AUTHOR_INPUT_NEEDED (pairwise test type) with Bonferroni correction across AUTHOR_INPUT_NEEDED (number of comparisons in the family); reported pairwise p values are AUTHOR_INPUT_NEEDED (Bonferroni-adjusted or unadjusted). The association between condition and task abandonment was tested with a chi-square test, and the association between self-reported effort and workload with AUTHOR_INPUT_NEEDED (correlation coefficient type). Accuracy was analysed with AUTHOR_INPUT_NEEDED. Participants were allocated to conditions by AUTHOR_INPUT_NEEDED.\n\nResults\nThe final sample comprised 90 participants (n = 30 per condition); two participants who failed the attention check were excluded before analysis. Perceived workload differed across the three feedback conditions, F(2, 87) = AUTHOR_INPUT_NEEDED, p = AUTHOR_INPUT_NEEDED, partial eta squared = .07. Participants in the delayed-feedback condition reported higher workload than those in the immediate condition (M = 5.1, SD = 1.2 vs M = 4.4, SD = 1.3), t(58) = 2.17, p = .034 (AUTHOR_INPUT_NEEDED: adjusted or unadjusted; Bonferroni-adjusted p AUTHOR_INPUT_NEEDED), d = 0.56, confidence interval AUTHOR_INPUT_NEEDED. Workload did not differ significantly between the immediate and no-feedback conditions, t(58) = 0.82, p = .42, effect size AUTHOR_INPUT_NEEDED. The delayed and no-feedback conditions AUTHOR_INPUT_NEEDED (comparison result). Condition was associated with task abandonment, chi2(2, N = 90) = 6.35, p = .042, effect size AUTHOR_INPUT_NEEDED. Self-reported effort correlated with workload, r(88) = .31, p = .003. Accuracy in the delayed condition did not differ significantly from AUTHOR_INPUT_NEEDED (comparison group), AUTHOR_INPUT_NEEDED (test statistic and df), p = .07."
},
"reviewer_risk": [
"If the corrected F test p value lies near .05, a reviewer may question how much ... the omnibus result can bear for the pairwise follow-ups.",
"If the line 7 p value is unadjusted, the delayed versus immediate difference may not remain significant after the stated Bonferroni correction, which would change the central pairwise claim.",
"A reviewer may ask whether assumptions of the ANOVA and t-tests were checked and how the accuracy outcome was defined and measured."
],
"author_input_needed": [
"Which of F, df or p at line 6 is correct, and what are the exact values from the R output? (Results, line 6)",
"Are the pairwise p values at lines 7 and 8 Bonferroni-adjusted or unadjusted, and how many comparisons formed the family? (Statistical analysis and Results)",
"What type of test was used for the pairwise comparisons, for example independent-samples t-tests or post hoc contrasts? (Statistical analysis)",
"What is the result of the delayed versus no-feedback comparison? (Results)",
"Which correlation coefficient was used for effort and workload? (Statistical analysis)",
"What test, comparison group, statistic and df produced the accuracy p = .07? (Statistical analysis and Results, line 11)",
"What are the effect sizes and confidence intervals for lines 7, 8 and 9? (Results)",
"How were participants allocated to the three feedback conditions? (Statistical analysis or Methods)",
"Was the ANOVA a one-way between-participants design? (Statistical analysis)"
],
"prescan_responses": [
{
"ref": "S1",
"verdict": "confirmed",
"note": "Line 6 reports F(2, 87) = 3.12, p = .012, but T1 gives an implied p of about 0.0491 (range 0.0489 to 0.0494); both are below .05, so this is a same-side inconsistency (I1)."
},
{
"ref": "S2",
"verdict": "confirmed",
"note": "Line 11 calls p = .07 marginally significant and infers that delayed feedback drives poorer performance (I3)."
},
{
"ref": "S3",
"verdict": "confirmed",
"note": "Participants are compared across three conditions (line 5) and the text never states how they were allocated; blinding is less central for a self-report study, but allocation should be stated (I6)."
}
]
}
Worked example: draft
The page's "Draft the Statistical analysis for the neuroscience paper" example (title
Kcnq5 spine density (Fig. 2-3), target nature). It is the neuroscience
Results section from the page's first example, plus three lines of design notes from the authors
(3 mice per genotype, 18 cells from the same mice, unpaired t-tests in GraphPad Prism, nothing
corrected, why two mice were excluded), plus the audit of that section, turned into text by
Recon.auditText exactly as the "Draft the section from this audit" button does. The
prescan found one test result, T1, t(4) = 2.10, P = 0.03, which implies
p = 0.1037 (range 0.1031 to 0.1042): a decision_error. It raised thirteen flags,
3 high (stat_decision_error, n_subsample_unit for spines and cells,
multiple_comparisons), 6 medium and 4 low, so
browser_verdict is not_ready. Because the target is Nature,
tails_unstated, df_missing and n_range are medium and
repeats_unstated appears. Its Idempotency-Key from the page is
stats-desk:draft:1cehxnmtrgdzo:a1.
The request body, with text, facts and audit abridged:
{
"task": "draft",
"target": "nature",
"journal": "",
"title": "Kcnq5 spine density (Fig. 2-3)",
"design": "Kcnq5 knockout and wild-type littermates, 3 mice per genotype for imaging (one hippocampal slice per mouse), 3 mice per genotype for behaviour.\nSpines were measured on CA1 pyramidal neurons, several dendrites per mouse. mEPSCs: 18 cells from the same 3 mice per genotype.\nAll tests were unpaired t-tests in GraphPad Prism; nothing was corrected for multiple comparisons. The two excluded mice froze during the test phase.",
"text": "1| Statistical analysis\n2| Data are presented as mean ± error bars. Statistical significance was assessed using unpaired... [abridged here: all 15 lines are shown decoded below]",
"facts": "{\"target\":\"nature\",\"browser_verdict\":\"not_ready\",\"flags\":[{\"id\":\"S1\",\"severity\":\"high\",\"priority\":\"P0\",\"category\":\"stat_... [abridged here: decoded below]",
"question": "",
"audit": "I1 [S1] P0 stat_decision_error at line 10: The reported P = 0.03 does not match the statistic: per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.... [abridged here: decoded below]"
}
Its text, decoded:
1| Statistical analysis
2| Data are presented as mean ± error bars. Statistical significance was assessed using unpaired t-tests or one-way ANOVA in GraphPad Prism. P < 0.05 was considered significant. *P < 0.05, **P < 0.01, ***P < 0.001.
3|
4| Results
5| Loss of Kcnq5 increased dendritic spine density in CA1 pyramidal neurons (n = 142 spines from 3 mice per genotype; P < 0.001, t-test; Fig. 2b).
6| Spine head volume was also larger in knockout neurons (n = 96 spines, P < 0.01; Fig. 2c), whereas spine length was not significantly different (ns; Fig. 2d).
7| Miniature EPSC frequency was higher in knockout slices (t = 3.41, P < 0.01; n = 18 cells), but mEPSC amplitude did not change (P > 0.05).
8| The increase in spine density was significant in males but not in females, indicating that the effect of Kcnq5 loss is sex-specific (Fig. 2e).
9| Across the five hippocampal subfields, knockout mice showed higher c-Fos counts in CA1, CA3 and DG (all P < 0.05) but not in CA2 or subiculum (Fig. 3a).
10| Novel object recognition was impaired in knockout mice (discrimination index 0.12 ± 0.05 vs 0.31 ± 0.04; t(4) = 2.10, P = 0.03; n = 3 mice per group).
11| Two outlier mice were excluded from the behavioural analysis.
12| Representative images are shown in Fig. 3b. These results prove that Kcnq5 controls spine formation.
13|
14| Figure 2 legend
15| b-d, Spine density, head volume and length; each dot is one spine. Error bars, mean ± s.e.m. n = 3-4 mice per genotype.
Its audit, decoded (abbreviated here to its first four lines; the full string is 6,417 characters in 30 lines):
I1 [S1] P0 stat_decision_error at line 10: The reported P = 0.03 does not match the statistic: per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.1042, two-tailed), which is above 0.05. Fix: Recheck the statistic, df and P value against the Prism output and report the correct values; if the test was one-tailed this must be stated and justified in advance, otherwise revise the claim on line 10.
I2 [S2] P0 n_subsample_unit at line 5: Spines (lines 5-6) and cells (line 7) are used as n for t tests, while the independent unit is the mouse (3 per genotype); nothing states averaging per animal or a nested model. Fix: Analyse per-mouse means (n = mice) or use a mixed model with mouse as a random effect, and report the number of mice, cells and spines per group for each panel.
I3 [S4] P0 significance_difference at line 8: A sex-specific effect is inferred from the effect being significant in one sex and not the other; no interaction test is reported. Line 9 similarly contrasts subfields by significance. Fix: Report a test of the genotype by sex interaction (or a direct comparison of the effects) with n per sex, or remove the sex-specific interpretation.
I4 [S3] P0 multiple_comparisons at line 9: Five subfields are compared and reported with threshold P values only, with no correction or declared family of planned comparisons; no multiple-comparison correction is named anywhere in the text. Fix: Name the test for line 9, define the family of comparisons and state the correction applied (or that none was applied and why), and give exact P values per subfield.
[... abridged here: 26 more lines, the remaining issues I5-I16 and the author questions ...]
Its facts, decoded:
{
"target": "nature",
"browser_verdict": "not_ready",
"flags": [
{
"id": "S1",
"severity": "high",
"priority": "P0",
"category": "stat_decision_error",
"message": "T1 (line 10): t(4) = 2.10, P = 0.03 implies p = 0.104 (two-tailed where applicable; range 0.103 to 0.104 allowing for rounding), so the reported p = 0.03 is on the wrong side of 0.05.",
"lines": [
10
]
},
{
"id": "S2",
"severity": "high",
"priority": "P0",
"category": "n_subsample_unit",
"message": "n counts spines, cells and nothing says they were averaged per independent unit or modelled as nested - possible pseudoreplication, which shrinks p values.",
"lines": [
5,
6,
7
]
},
{
"id": "S3",
"severity": "high",
"priority": "P0",
"category": "multiple_comparisons",
"message": "6 reported p value(s) or test result(s), but no multiple-comparison correction or declared family of planned comparisons is named.",
"lines": [
10
]
},
{
"id": "S4",
"severity": "medium",
"priority": "P1",
"category": "significance_difference",
"message": "One effect is significant and another is not, and the text reads that as the two effects differing. A difference in significance is not a significant difference - test the interaction or compare the effects directly.",
"lines": [
6,
7,
8,
9
]
},
{
"id": "S5",
"severity": "medium",
"priority": "P1",
"category": "tails_unstated",
"message": "The text never says whether tests were one- or two-tailed - Nature's statistics section must state it.",
"lines": []
},
{
"id": "S6",
"severity": "medium",
"priority": "P1",
"category": "df_missing",
"message": "1 t or F statistic(s) printed without degrees of freedom; Nature asks for t and F values with their degrees of freedom.",
"lines": [
7
]
},
{
"id": "S7",
"severity": "medium",
"priority": "P1",
"category": "n_range",
"message": "n is given as a range rather than exact values. Give the exact n for each group or panel, as Nature requires.",
"lines": [
15
]
},
{
"id": "S8",
"severity": "medium",
"priority": "P1",
"category": "exclusions_vague",
"message": "Exclusions or outliers are mentioned without the rule, when it was set, or how many were removed.",
"lines": [
11
]
},
{
"id": "S9",
"severity": "medium",
"priority": "P1",
"category": "repeats_unstated",
"message": "Representative results are shown without saying how many times the experiment was repeated - Nature asks for the repeat count.",
"lines": [
12
]
},
{
"id": "S10",
"severity": "low",
"priority": "P2",
"category": "ns_without_p",
"message": "A result is called not significant (or ns) without its p value or effect estimate. Absence of significance is not evidence of no effect.",
"lines": [
6
]
},
{
"id": "S11",
"severity": "low",
"priority": "P2",
"category": "overclaim",
"message": "Strength words (\"highly significant\", \"proves\", \"conclusively\") lean on the p value rather than the effect size and its uncertainty.",
"lines": [
12
]
},
{
"id": "S12",
"severity": "low",
"priority": "P2",
"category": "sem_used",
"message": "s.e.m. describes the precision of a mean, not the spread of the data; with small n it makes groups look tighter. Pair it with n or show s.d. or a CI.",
"lines": [
15
]
},
{
"id": "S13",
"severity": "low",
"priority": "P2",
"category": "randomization_blinding",
"message": "Animals or participants are compared, and the text says nothing about randomization or blinding.",
"lines": []
}
],
"stats": [
{
"id": "T1",
"line": 10,
"test": "t",
"df1": 4,
"df2": null,
"value": 2.1,
"reported": "t(4) = 2.10, P = 0.03",
"p_reported": "p = 0.03",
"p_computed": 0.1037,
"p_range": [
0.1031,
0.1042
],
"status": "decision_error"
}
],
"n_mentions": [
{
"line": 5,
"text": "n = 142 spines from 3 mice per genotype",
"unit": "subsample",
"unit_word": "spines",
"range": false
},
{
"line": 6,
"text": "n = 96 spines, P < 0.01; Fig. 2",
"unit": "subsample",
"unit_word": "spines",
"range": false
},
{
"line": 7,
"text": "n = 18 cells), but mEPSC amplit",
"unit": "subsample",
"unit_word": "cells",
"range": false
},
{
"line": 10,
"text": "n = 3 mice per group).",
"unit": "independent",
"unit_word": "mice",
"range": false
},
{
"line": 15,
"text": "n = 3-4 mice per genotype.",
"unit": "independent",
"unit_word": "mice",
"range": true
}
],
"tests_named": [
"t-test",
"one-way ANOVA",
"ANOVA"
],
"corrections_named": [],
"software_named": [
"GraphPad Prism"
],
"summary_conventions": [
"mean ± s.e.m."
],
"named_in_design_notes": {
"tests": [
"t-test"
],
"corrections": [],
"software": [
"GraphPad Prism"
]
},
"tails_stated": false,
"one_tailed_declared": false,
"p_counts": {
"exact": 1,
"threshold": 5,
"ns": 2
},
"clipped": null
}
The reply (parsed; shown in full). The drafted section keeps every fact the authors gave: the mouse
as the independent unit, 3 mice per genotype, the 142 and 96 spines and 18 cells as reported,
unpaired t-tests in GraphPad Prism, no correction, the freezing exclusion. Everything they did not
give is an AUTHOR_INPUT_NEEDED: the Prism version, the tails, how spines and cells were
analysed relative to the mouse, where one-way ANOVA was used. The audit recommended per-mouse means or
a mixed model; the draft does not claim either was done. Four P0 issues remain open (the
decision error on line 10, spines and cells as n, the sex-specific claim, the uncorrected subfields),
so the drafted section is not_ready. All thirteen flags are answered.
check-reply.js reports 0 disagreements (every number and method in the draft and the
legend template comes from the text or the design notes) and one item to review: 20
placeholders against 15 author questions.
{
"lane": "draft",
"verdict": "not_ready",
"headline": "The drafted section records the supplied design, software and the absence of multiplicity correction, but the central spine and mEPSC results were analysed with spines and cells as n, and the novel object recognition P value is inconsistent with its statistic, so it cannot yet be finalised.",
"tldr": [
"The draft states the mouse as the independent unit, 3 mice per genotype, unpaired t-tests in GraphPad Prism, no multiple-comparison correction and the freezing-based exclusion, all from the design notes.",
"The spine (lines 5-6) and mEPSC (line 7) results use spines and cells as n; how they relate to the 3 mice per genotype must be resolved before the section is usable.",
"Per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.1042), not the reported P = 0.03 on line 10; the value must be rechecked against the Prism output.",
"The sex-specific claim (line 8) rests on a difference in significance and the five-subfield c-Fos comparison (line 9) has no correction or family definition.",
"Nature asks for the tails of each test, exact n values, exact P values, t values with degrees of freedom and the number of repeats for representative images; these remain placeholders."
],
"issues": [
{
"id": "I1",
"ref": "S1",
"priority": "P0",
"category": "stat_decision_error",
"line": 10,
"evidence": "t(4) = 2.10, P = 0.03",
"problem": "Per T1, t(4) = 2.10 implies p_computed 0.1037 (range 0.1031 to 0.1042, two-tailed), on the other side of 0.05 from the reported P = 0.03.",
"why": "A reviewer will question whether the novel object recognition impairment is supported and whether an undeclared one-tailed test was used.",
"fix": "Recheck the statistic, df and P value against the Prism output and report the correct values; if the claim no longer holds at P < 0.05, revise the wording on line 10."
},
{
"id": "I2",
"ref": "S2",
"priority": "P0",
"category": "n_subsample_unit",
"line": 5,
"evidence": "n = 142 spines from 3 mice per genotype; P < 0.001, t-test",
"problem": "Spines (lines 5-6) and cells (line 7) are used as n for t-tests, while the design notes give 3 mice per genotype as the source of all spines and all 18 cells.",
"why": "Treating subsamples as independent inflates n and shrinks P values; with 3 mice per genotype the central claim may not be supported.",
"fix": "State whether the analysis used per-mouse means (n = mice) or a model accounting for mouse; if neither, reanalyse and report mice, cells and spines per group for each panel."
},
{
"id": "I3",
"ref": "S4",
"priority": "P0",
"category": "significance_difference",
"line": 8,
"evidence": "The increase in spine density was significant in males but not in females, indicating that the effect of Kcnq5 loss is sex-specific",
"problem": "A sex-specific effect is inferred from significance in one sex and not the other; no interaction test or n per sex is given.",
"why": "A difference in significance is not a significant difference; a reviewer will ask for the genotype by sex interaction.",
"fix": "Report a test of the genotype by sex interaction with n per sex, or remove the sex-specific interpretation."
},
{
"id": "I4",
"ref": "S3",
"priority": "P0",
"category": "multiple_comparisons",
"line": 9,
"evidence": "knockout mice showed higher c-Fos counts in CA1, CA3 and DG (all P < 0.05) but not in CA2 or subiculum",
"problem": "Five subfields are compared with threshold P values; the design notes confirm no correction was applied and no family of planned comparisons is defined.",
"why": "Uncorrected multiple comparisons raise the chance of false positives, and the subfield pattern is read from differences in significance.",
"fix": "Define the family of comparisons and either apply a correction or state why none was applied, and report exact P values per subfield."
},
{
"id": "I5",
"ref": "",
"priority": "P1",
"category": "test_conflict",
"line": 2,
"evidence": "Statistical significance was assessed using unpaired t-tests or one-way ANOVA in GraphPad Prism.",
"problem": "Line 2 names one-way ANOVA, while the design notes say all tests were unpaired t-tests.",
"why": "A reviewer cannot tell which comparisons used which test, and Nature asks for F values and degrees of freedom for ANOVAs.",
"fix": "Confirm whether one-way ANOVA was used; if so name the comparison and report F(df1, df2), otherwise remove it from line 2."
},
{
"id": "I6",
"ref": "S8",
"priority": "P1",
"category": "exclusions_vague",
"line": 11,
"evidence": "Two outlier mice were excluded from the behavioural analysis.",
"problem": "The design notes give the reason (freezing during the test phase), but not whether the rule was prespecified, which genotype the mice came from, or the result with them included.",
"why": "Undisclosed timing of an exclusion rule in a 3-per-group comparison could change the result.",
"fix": "State the prespecified criterion, the genotype of each excluded mouse, and whether conclusions hold when they are included; describe them as excluded for freezing rather than as outliers."
},
{
"id": "I7",
"ref": "S5",
"priority": "P1",
"category": "tails_unstated",
"line": 0,
"evidence": "",
"problem": "Neither the text nor the design notes say whether tests were one- or two-tailed.",
"why": "Nature asks for the Methods statistics section to state whether tests were one- or two-tailed.",
"fix": "Add the tails for all tests to the Statistical analysis."
},
{
"id": "I8",
"ref": "S7",
"priority": "P1",
"category": "n_range",
"line": 15,
"evidence": "n = 3-4 mice per genotype.",
"problem": "The legend gives a range, which conflicts with line 5 and the design notes (3 mice per genotype).",
"why": "Nature asks for exact n values, individual values rather than a range if n varied.",
"fix": "Give the exact number of mice per genotype for each of panels b-d and reconcile with line 5."
},
{
"id": "I9",
"ref": "S6",
"priority": "P1",
"category": "df_missing",
"line": 7,
"evidence": "t = 3.41, P < 0.01; n = 18 cells",
"problem": "The t value has no degrees of freedom and the P value is a threshold.",
"why": "Nature asks for t values and degrees of freedom for t-tests.",
"fix": "Report t(df) with the exact P value and the cells and mice per genotype."
},
{
"id": "I10",
"ref": "",
"priority": "P1",
"category": "threshold_p_values",
"line": 6,
"evidence": "n = 96 spines, P < 0.01; Fig. 2c",
"problem": "All results except line 10 give threshold P values or ns, with no test statistics, effect sizes or intervals.",
"why": "Nature asks for exact P values for significant and non-significant results where relevant; threshold values hide effect size and uncertainty.",
"fix": "Report group means with the defined dispersion, the test statistic with df, and the exact P value for each comparison."
},
{
"id": "I11",
"ref": "S9",
"priority": "P1",
"category": "repeats_unstated",
"line": 12,
"evidence": "Representative images are shown in Fig. 3b.",
"problem": "The number of times the representative measurement was repeated is not given.",
"why": "Nature asks for the number of times representative measurements were repeated.",
"fix": "State in the Fig. 3 legend how many independent experiments or mice the representative images reflect."
},
{
"id": "I12",
"ref": "",
"priority": "P1",
"category": "summary_convention_undefined",
"line": 10,
"evidence": "discrimination index 0.12 ± 0.05 vs 0.31 ± 0.04",
"problem": "Line 2 says mean ± error bars without defining them; only the Fig. 2 legend says s.e.m., and the ± values on line 10 are undefined.",
"why": "Nature asks for error bars to be defined throughout the figures.",
"fix": "State what the ± values and error bars represent for every panel and Results value."
},
{
"id": "I13",
"ref": "S13",
"priority": "P2",
"category": "randomization_blinding",
"line": 0,
"evidence": "",
"problem": "Randomization and blinding are not mentioned in the text or design notes.",
"why": "Unblinded spine counting and behavioural scoring are a common reviewer concern.",
"fix": "State whether allocation, spine quantification, electrophysiology and behavioural scoring were randomized or blinded."
}
],
"draft_text": "Statistical analysis. All statistical tests were performed in GraphPad Prism (version AUTHOR_INPUT_NEEDED). Spine density, head volume and length (Fig. 2b-d) are presented as mean ± s.e.m.; the summary convention for other panels, including the discrimination index values, is AUTHOR_INPUT_NEEDED, as is the n over which each s.e.m. was computed. Kcnq5 knockout mice and wild-type littermates were compared, and the independent experimental unit was the mouse. For spine imaging, 3 mice per genotype were analysed (one hippocampal slice per mouse); spines were measured on several dendrites of CA1 pyramidal neurons per mouse (142 spines for spine density and 96 spines for head volume; spines per genotype and per mouse: AUTHOR_INPUT_NEEDED; exact mice per genotype for Fig. 2b-d: AUTHOR_INPUT_NEEDED). mEPSCs were recorded from 18 cells from the same 3 mice per genotype (cells per genotype and per mouse: AUTHOR_INPUT_NEEDED). Behaviour was tested in 3 mice per genotype (number of mice tested before exclusion: AUTHOR_INPUT_NEEDED). Spines and cells are subsamples within each mouse; they were analysed as AUTHOR_INPUT_NEEDED (per-mouse means with n = mice, or a model that accounts for mouse). Comparisons between genotypes used unpaired t-tests (one- or two-tailed: AUTHOR_INPUT_NEEDED). Use of one-way ANOVA, the comparison it was applied to, and its F value and degrees of freedom: AUTHOR_INPUT_NEEDED. Test used to compare the genotype effect between males and females (Fig. 2e), with n per sex: AUTHOR_INPUT_NEEDED. No correction for multiple comparisons was applied, including across the five hippocampal subfields for c-Fos counts (Fig. 3a); family of planned comparisons and rationale: AUTHOR_INPUT_NEEDED. P < 0.05 was considered significant. Test statistics are reported with degrees of freedom and exact P values for all comparisons, including non-significant ones (values: AUTHOR_INPUT_NEEDED). Two mice were excluded from the behavioural analysis because they froze during the test phase; whether this criterion was set before analysis, the genotype of each excluded mouse and the result with them included: AUTHOR_INPUT_NEEDED. Representative images (Fig. 3b) reflect AUTHOR_INPUT_NEEDED independent experiments. Randomization of allocation and blinding during spine quantification, electrophysiology and behavioural scoring: AUTHOR_INPUT_NEEDED.",
"reporting_notes": {
"n_definition": "Mouse as the independent unit; 3 mice per genotype for imaging with one slice per mouse, 18 mEPSC cells from the same mice, and 3 mice per genotype for behaviour, from the design notes; 142 and 96 spines from lines 5-6; per-mouse breakdown left as placeholders.",
"tests_models": "Unpaired t-tests from the design notes and line 2; the one-way ANOVA on line 2 conflicts with the design notes and is left as a placeholder, as is the analysis of spines and cells relative to mouse and the sex comparison.",
"multiple_comparisons": "States that no correction was applied, from the design notes; the family of planned comparisons and rationale are a placeholder.",
"software": "GraphPad Prism from line 2 and the design notes; version is a placeholder.",
"tails": "Not stated in text or design notes; left as a placeholder.",
"unresolved": [
"GraphPad Prism version",
"Summary convention for panels other than Fig. 2b-d, including the ± values of the discrimination index, and the n over which s.e.m. was computed",
"Spines per genotype and per mouse for spine density and head volume",
"Exact number of mice per genotype for Fig. 2b-d (line 5 gives 3, the legend gives 3-4)",
"mEPSC cells per genotype and per mouse",
"Number of behaviour mice tested before exclusion",
"How spines and cells were analysed relative to mouse (per-mouse means or a nested model)",
"Whether tests were one- or two-tailed",
"Whether and where one-way ANOVA was used, with F and degrees of freedom",
"Test for the genotype by sex comparison and n per sex",
"Family of planned comparisons and rationale for no correction",
"Exact test statistics, degrees of freedom and P values for each comparison",
"Whether the freezing exclusion was prespecified, genotype of each excluded mouse, and result with them included",
"Number of independent experiments behind the representative images in Fig. 3b",
"Randomization and blinding"
]
},
"legend_template": "Each dot represents one AUTHOR_INPUT_NEEDED (mouse, cell or spine); bars show mean ± s.e.m. of n = AUTHOR_INPUT_NEEDED mice per genotype (AUTHOR_INPUT_NEEDED spines or cells per mouse); AUTHOR_INPUT_NEEDED (one- or two-tailed) unpaired t-test, no correction for multiple comparisons; exact P values: AUTHOR_INPUT_NEEDED.",
"author_input_needed": [
"Which GraphPad Prism version was used? (Statistical analysis)",
"Were spine and cell measurements averaged per mouse or analysed with a model accounting for mouse, and with what model? (Statistical analysis)",
"How many spines per genotype and per mouse contributed to spine density and head volume? (Statistical analysis; Results, lines 5-6)",
"What are the exact numbers of mice per genotype for Fig. 2b-d, given line 5 says 3 and the legend says 3-4? (Figure 2 legend)",
"How many of the 18 mEPSC cells came from each genotype and each mouse, and what is the df for t = 3.41? (Statistical analysis; Results, line 7)",
"Were tests one- or two-tailed? (Statistical analysis)",
"Was one-way ANOVA used, for which comparison, and what are F and its degrees of freedom? (Statistical analysis; Results)",
"What are the correct t statistic, df and exact P value for the novel object recognition comparison? (Results, line 10)",
"Was a genotype by sex interaction tested, and what are the n per sex? (Statistical analysis; Results, line 8)",
"What is the family of planned comparisons for the five c-Fos subfields and why was no correction applied? (Statistical analysis; Results, line 9)",
"What are the exact P values and test statistics for all comparisons, including spine length and mEPSC amplitude? (Results)",
"What do the ± values on line 10 and the error bars in panels other than Fig. 2b-d represent? (Statistical analysis; legends)",
"How many behaviour mice were tested before exclusion, was the freezing criterion set before analysis, which genotype were the two excluded mice, and does the result change with them included? (Statistical analysis; Results, line 11)",
"How many times were the representative measurements in Fig. 3b repeated? (Figure 3 legend)",
"Were allocation, spine quantification, electrophysiology and behavioural scoring randomized or blinded? (Statistical analysis)"
],
"prescan_responses": [
{
"ref": "S1",
"verdict": "confirmed",
"note": "Line 10 reports t(4) = 2.10, P = 0.03; per T1 the statistic implies p_computed 0.1037 (range 0.1031 to 0.1042), above 0.05."
},
{
"ref": "S2",
"verdict": "confirmed",
"note": "Lines 5-7 use spines and cells as n; the design notes say all came from 3 mice per genotype and describe no per-mouse averaging or nested model."
},
{
"ref": "S3",
"verdict": "confirmed",
"note": "Line 9 compares five subfields with threshold P values and the design notes confirm nothing was corrected for multiple comparisons; no family is defined."
},
{
"ref": "S4",
"verdict": "confirmed",
"note": "Line 8 infers a sex-specific effect from significance in males but not females without an interaction test; line 9 contrasts subfields the same way."
},
{
"ref": "S5",
"verdict": "confirmed",
"note": "Neither the text nor the design notes state whether tests were one- or two-tailed."
},
{
"ref": "S6",
"verdict": "confirmed",
"note": "Line 7 gives t = 3.41 without degrees of freedom."
},
{
"ref": "S7",
"verdict": "confirmed",
"note": "Line 15 gives n = 3-4 mice per genotype, a range that also conflicts with line 5 and the design notes (3 per genotype)."
},
{
"ref": "S8",
"verdict": "confirmed",
"note": "Line 11 gives no rule; the design notes add that the mice froze during the test phase, but not whether the rule was prespecified or their genotype."
},
{
"ref": "S9",
"verdict": "confirmed",
"note": "Line 12 shows representative images in Fig. 3b without the number of repeats."
},
{
"ref": "S10",
"verdict": "confirmed",
"note": "Line 6 reports spine length as ns and line 7 mEPSC amplitude as P > 0.05, without exact P values or estimates."
},
{
"ref": "S11",
"verdict": "confirmed",
"note": "Line 12 says these results prove that Kcnq5 controls spine formation, beyond density data from 3 mice per genotype."
},
{
"ref": "S12",
"verdict": "confirmed",
"note": "Line 15 uses s.e.m. with each dot one spine and 3-4 mice per genotype, so the n behind the s.e.m. is unclear."
},
{
"ref": "S13",
"verdict": "confirmed",
"note": "Mice are compared and neither the text nor the design notes mention randomization or blinding."
}
]
}
Truncation and partial results
When the balance sits between min_credits and hold_credits, the run is not
refused: it executes with a reduced output cap and reports truncated: true, in the
done event of /run-stream and on the job from GET /jobs/{id}.
What you hold is then a prefix of the reply, and it will not parse as it stands. The web page closes
the cut-off JSON (Recon.closeJson in recon.js: close an open string, drop
a dangling comma, give a dangling key null, close every open array and object, and if
that still does not parse, cut back to the previous comma and try again), parses what is left, and
shows the sections that arrived as "N of M sections recovered", out of the lane's
11 keys (10 for draft). A stream that ends early with an error gets the same treatment
on the deltas received so far.
The keys arrive in contract order, so a cut usually costs the tail: prescan_responses
and author_input_needed go first, then the lane body. A truncated audit can therefore
look complete while flags are unanswered, and a cut inside revision.text or
draft_text leaves you half a section with placeholders that no question backs. Never
paste a truncated section: check the flag, top up, and resubmit with the attempt suffix on the
Idempotency-Key incremented (stats-desk:audit:<hash>:a2).
If a complete reply will not parse as one JSON object, the page retries once, as the next attempt,
with a retry_note saying what was wrong: "Your previous reply was not the single valid
JSON object the instructions require (the parse error). Reply again with ONLY the JSON
object for task 'lane' - no prose, no code fences; every key present (empty arrays where
there is nothing to say)." Do the same: keep the hash, bump the attempt, add
retry_note. It never retries a truncated reply that way; that one needs credits, not a
reformat. If the retry still does not parse, the page shows the raw reply.