← 返回 Skill 列表
extension
分类: 开发与工程API Key 暂未确认

同行审计核验

核验他人/他模型交付的工作产物审计报告:把"已完成""已验证"等自述逐条拆成可证伪声明并独立重算,而不是采信叙述。 核心能力: 雪崩判据识别伪造哈希——声称值与实测值共享 ≥8 位十六进制前缀后分叉,即"真前缀+编造尾巴"的构造值,而非过期摘要;12 位相同 ≈ 2⁻⁴⁸,不可解释为巧合。 双工具交叉重算每个摘要(hashlib / sha256sum / Get-FileHash),不一致即停。 分层取证:已声明 → 已注册 → 已派发 owner → 新会话暴露 → 实际执行 → 有用。证据属于某一层永远不能证明下一层;连接测试通过 ≠ 工具进了模型目录;新增策略层若被旧分支遮蔽,纯函数单测 100% 全绿也照样是真实绕过。 测试鉴别力:先在修复前状态证明新测试能红(假绿比红测试更危险,因其会被下游信任);落点必须在消费者层。 测试报告自己的验证命令——命令 key 在不存在的字段上会对每一行静默返回错值,而散文结论仍成立。 随机后端只引用带 n 的区间;重跑随机生成器而非引用其数字。 闭环声明("已修复""已轮换密钥")同样是自述,未亲自实测不算关闭。 秘密卫生:只发布布尔/计数/指纹。按 mtime 排序区分"死残留"与"活的二次喷发源"——后者说明重定源在 shell 之外(网关快照、父进程重注入),清洗无效,必须轮换。 交付门:指定输出文件即交付物本身,先建完整骨架(未证处写 INCONCLUSIVE),预算耗尽而无落盘产物 = 工作流失败。 含作者侧报告骨架(反向:写一份供同行攻击的报告)、审阅回执模板(判词表+哈希台账+评分带+机读 JSON 块)、两个可重跑探针脚本。适用于代码、配置、插件注册、会话产物等多类可证伪交付物。

person作者: dreamFTYhubModelScope

Peer Audit Verification (Minimal Router)

Independently verify reports that make falsifiable claims about a machine or codebase: SHA-256 matrices, file states/line counts, test-suite results, CLI availability, cross-tier sync, session provenance. The report's own "verification suite" is the first thing to execute — a report that fails it is the headline finding.

Resources (Progressive Disclosure):

  • Long-task completion gate, artifact-first discipline, placeholder residue checks: references/completion-gate.md
  • Layered evidence (declared → registered → dispatched → exposed → executed → useful), probing mechanism instead of narrative, dead axes and legacy-branch shadowing, closure claims are claims too: references/layered-verification.md
  • Test power: prove it can go red first, false-green/false-red diagnosis, the assertion must land on the consumer layer: references/test-power.md
  • Command recipes (hash cross-check, PowerShell-from-bash, JSON slicing, ledger authoring traps, stochastic-backend measurement): references/command-recipes.md
  • Reply form to hand a reviewer model (verdict vocabulary, hash ledger, scoring rubric, machine-readable block): templates/audit-reply-form.md
  • Re-runnable probe: scripts/verify_hashes.py — feed it a JSON list of {label, path, claimed} and it recomputes + flags forged tails
  • Re-runnable probe: scripts/secret_residue_scan.py — find a credential's plaintext residue by mtime-sorted boolean scan, distinguishing a dead leftover from a live re-emitter; never prints the secret
  • Author-side skeleton when you are writing the report for peer models to attack: templates/llm-audit-report.md — section contract, claim-table format, trap-list format, no-terminal track

0. Completion gate — the reply file is the deliverable

When the requested result is a report, audit reply, review form, or machine-readable handoff, treat the named output file as the primary deliverable, not a summary to append after the investigation. Create the complete skeleton (every required section, table, signature, machine-readable block) immediately after preflight, using INCONCLUSIVE wherever evidence is not yet available — never omit a required field.

A forced budget stop with no artifact on disk is a workflow failure, even if the findings are strong. Budget allocation, resumption, and the mechanical completeness checks: references/completion-gate.md.

1. Verification order

  1. Enumerate falsifiable claims; prioritize published numerics (hashes, counts) → file states → CLI behavior → content claims → session narrative.
  2. Execute the report's own verification commands literally (translate PowerShell → git-bash as needed).
  3. Recompute every digest from disk with ≥2 independent tools — Python hashlib, coreutils sha256sum, PowerShell Get-FileHash; all must agree before you assert any value.
  4. Run each claimed-vs-computed pair through the avalanche rule (§2). Never repeat a claimed digest in your verdict without having reproduced it.
  5. Run test suites with the project venv; verify the declared total (Ran N tests + trailing OK, not just exit code) and that the named tests physically exist in the file.
  6. CLI claims: command -v, --version, read the shim/launcher chain end-to-end; check the USER PATH registry, not only the session-inherited PATH (they can differ).
  7. Spot-read the exact sections/links that content claims cite; verify referenced files exist.
  8. Deliver verdicts as claim → method → actual → verdict tables; keep substantive conclusions and published-value verdicts separate (§3).
  9. Recompute the reviewer's own arithmetic — weighted totals, percentages, rates — with a tool, exactly as you would recompute their hashes. An error that crosses the reviewer's own band boundary (reported total falls in "verified" while the true total falls in "fixes needed") is a material finding, not a nitpick: state the computed value, the delta, and the band it actually lands in.
  10. Check reply completeness mechanically: count unfilled placeholder/required-field markers and validate that any embedded JSON block parses. Cheap, and it is the reviewer's own §0 obligation — don't adopt a reply as complete without it.

Four rules that need their own references (each has a documented failure signature):

  • Prove each test can fail before believing it — false green is worse than a red test because it is trusted downstream; also the symmetric trap of the false red you author yourself. → references/test-power.md
  • Treat the report's own verification command as a claim too — a command keying on a nonexistent field silently returns the wrong value for every row while the prose conclusion still holds; and a claim table disagreeing with its own prose is a finding. → references/layered-verification.md
  • Re-run stochastic generators instead of quoting their numbers — a byte-identical artifact across two independent runs proves the numbers were not cherry-picked; a changed hash proves the claim is a single observation. → references/layered-verification.md
  • A closure report is a claim, whoever it comes from — "fixed", "I rotated the key", "resolved" are all self-reports, graded exactly as the original artifact. A blocker is not closed until you measured it. → references/layered-verification.md

2. Avalanche rule — forged / mis-copied digest detection

SHA-256 avalanches: one input-bit flip changes ~50% of output bits. Genuine digests of differing content diverge from the FIRST nibble — they never share long prefixes. Therefore:

  • Claimed and computed digests share a long hex prefix (≥8–12 chars) then diverge → the claimed value was constructed (real prefix + invented tail — the "saw a truncated display, completed it plausibly" pattern), not produced by any computation. 12 shared chars = 48 bits ≈ 2⁻⁴⁸ per pair; across several pairs, impossible.
  • Distinguish fabrication from staleness: a stale-but-genuine digest (file edited after hashing) diverges everywhere, first nibble included. Prefix-sharing ⇒ constructed, never "nearly matching".
  • Search the source tree for the TRUE tail fragments; if the real values appear nowhere, no genuine computation was ever recorded — state that.
  • Report it plainly as "not reproducible / constructed"; do not soften to "slightly different".

3. Parity ≠ values — label vs substance

"Cross-tier parity verified 100%" is two claims: (a) the copies are byte-identical (test copy-vs-copy directly) and (b) the published digest values equal reality (test against recomputed values). (a) can pass while (b) fails. Verdict them separately — apply the same split to any "VERIFIED" label: check the substance the label stands on.

3.1 Effective behavior beats diagnostics and declarations

Audit plugin/tool systems as separate layers: declared → registered → dispatched owner → exposed in a fresh session → executed → useful. Evidence from one layer never proves the next. A manifest warning is not a semantic verdict; a default (shadow) mode is not a health check; count by stable unique identifiers from machine-readable inventory, not human tables or substring matches; a connection test passing (mcp test → ✓ Connected) does not put tools in the model's catalog.

For framing adapters, distinguish byte-preserving, JSON-semantic-preserving, and schema-preserving — never call a json.loads/json.dumps bridge "verbatim forwarding". Probe mechanism claims against source, not narrative. A new spec-derived tier is only real if the consumer reads it; the sharper failure is a legacy branch that shadows it, which survives a 100%-green suite because unit tests only assert the pure function.

All of these — with the concrete grep/branch-order/A-B recipes and the "hold the reviewer to the same standard" rule: references/layered-verification.md.

4. Provenance of session claims

Session and artifact stores are per-product and per-agent-host. Before concluding "not found":

  • Calibrate the search first. A 0-result query is not proof of absence until you have shown the same query returns a term you know exists (browse with no arguments, or search a known-present string).
  • Match the id format to the product. Some hosts use timestamped session ids (YYYYMMDD_HHMMSS_xxxxxx), others use UUIDs for plan/cascade-style artifacts. A timestamp-shaped query will silently return nothing against a UUID-keyed store, and vice versa.
  • Locate each product's artifact directory before searching it. IDE-style hosts commonly keep generated plans, walkthroughs and audit reports under a per-user agent directory (e.g. ~/.<agent>/<host>/brain/<artifact-id>/), with an implementation plan, a walkthrough, a work-product audit report and a scratch subdirectory. Their session registry / pane layout usually lives in a separate single-file JSON under the app's roaming config directory. Find the directory first; do not assume one product's store holds another's artifacts.
  • A CLI and its IDE sibling may use entirely separate stores. Check which product a claim names before hunting, and never merge conclusions across two stores that do not share ids.
  • Giant single-line JSON needs a byte slice, not grep. Session registries and layout files are often one enormous line; grep -C context mode is useless there. Use a short Python snippet that reads the file and prints the byte window around the id, with errors="replace".

5. Environment pitfalls (Windows / git-bash)

  • Native tools get no MSYS path translation — pass C:/... forward-slash native paths, never git -C /c/... style MSYS paths.
  • PowerShell from bash: powershell.exe -NoProfile -Command '<script, double quotes inside>' — single-quote the whole script in bash; no temp .ps1 needed.
  • Scratch output for native tools: $LOCALAPPDATA/Temp (or the platform's temp dir), not /tmp.
  • Python <3.12: no backslashes inside f-string expressions — precompute; mind $-expansion in double-quoted bash strings.
  • wc -l counts newline bytes; a file without a trailing newline reports one fewer than its editor "logical lines" — state the convention whenever a line-count claim is disputed.
  • Probe targeted candidate subdirs when hunting files/strings across a config or home directory — broad root scans are slow and can fail outright.
  • Call the host agent's live skill/tool registry API before validating "skill X exists / was removed" claims — the in-context list is a session-start snapshot that can lag disk. Likewise, rendered CLI tables truncate cell values (e.g. some-long-plugin-name…) and wrap rows, so a full-name grep silently misses them and visual row counts miscount — take counts from machine-readable output (--json) and reconcile against what the rendered view appeared to say.
  • Probe a secret's presence with [ -n "$VAR" ]. Never use ${VAR:-fallback} or ${VAR:+yes} for that job: both expand and print the value when the variable is set, so a "did the key leak" probe becomes the leak. Announce key state as booleans, counts, lengths or a truncated fingerprint only.

6. Verdict delivery

  • Lead with the verdict. Then claim → method → actual tables; then the defect mechanism (with probability reasoning for forgery claims); then a paste-ready corrected block (e.g. a fixed hash-matrix JSON); then 1–3 short, concrete remediation offers.
  • Match the reader's language (Chinese request → Chinese delivery), and keep identifiers, paths, and raw error text untranslated.
  • Always deliver corrected values plus the located patch-target paths; never end at "mismatch — investigate further".
  • Do not unilaterally edit another agent's or another team's artifact: offer the erratum patch and the corrected block, and let the owner confirm.
  • When the user asks for a handoff/remediation prompt to return to the audited author, deliver severity-ordered conditions (blocker → major → minor), each with the repair direction and the observed-vs-expected pair, then an audit-boundary statement: files you modified and restored (with the hash that proves restoration), traps you hit, your own harness mistakes, and an explicit line that every number came from your own runs. Paste-ready, no prose padding.

7. Authoring an audit report for peer LLM review (mirror direction)

When the user asks for a "work product audit report" for other models to review, the deliverable is a falsification harness, not a narrative of what was done. Section contract: templates/llm-audit-report.md.

  • Every claim atomic and falsifiable: claim → the exact command that tests it → the value you observed → what result would refute it. A claim no command can refute is decoration; drop it or label it [UNVERIFIED].
  • The reviewer trap list is the highest-value section you will write. Enumerate every environment gate that flips a verdict — env-var gates, locale/codec, PATH/PYTHONPATH inheritance, proxies, working directory, line endings, lazy registration — each with the wrong conclusion it produces. On a healthy work product most bad verdicts originate in the reviewer's environment, so this section is what makes the report usable rather than merely readable.
  • Run your own protocol verbatim before delivering. The reproduction commands are the deliverable's load-bearing surface: a stale expected count, a command that only works because of your shell's exports, or an "expected" output that reads like a crash will each be found by the reviewer and cost the report its credibility. Paste the fast track into a clean shell and execute it exactly as written.
  • Label deliberate non-actions as decisions, not omissions, each with a reason and an explicit "a reviewer may reasonably disagree" — otherwise reviewers will 'fix' privacy gates and safety interlocks you intentionally left closed.
  • Diff a component against its vendor's own reference pattern before changing it, and say so when it already conforms. Third-party plugins often ship a faithful port of a published recipe; when the line-by-line comparison matches, the correct output is "no change, here is the mapping" — churn without evidence is a regression risk dressed as diligence, and a report that lists a component as "changed" implies a defect that may not exist. Reserve edits for a reproduced defect, and pin the legacy tier's values with a regression test when a spec-derived tier is added alongside them, so "behaviour unchanged" is provable rather than asserted.
  • State the non-claims explicitly: name what you did not verify and why (layers that only close in a fresh session, local patches that upstream updates will revert, uncharacterised quality), so absence of evidence is never read as a pass.
  • Carry a prior-round disposition table when an earlier review exists — every finding marked fixed / rejected-with-reason / still-open. Mark rejected items as rejected so the next reviewer does not inherit them as premises; correct the reviewer's mechanism, not just its verdict. When you accept a finding, record your own re-measurement in the errata row rather than the reviewer's figure: reviewers get mechanisms right and counts wrong often enough that the errata table is where the two must be reconciled, and it is also where a reviewer can see that a disputed finding was tested rather than argued. Fold accepted corrections into the reviewed artifact itself — an errata section that lives only in a reply file leaves the next reviewer reading the uncorrected claim.
  • Secret hygiene: publish booleans, counts and fingerprints, never key material, raw secret-bearing config, or auth headers. A hash ledger plus a behaviour check (HTTP 200 from the live endpoint) proves the same credential is in use without disclosing it. State the scan you ran (roots + tool), because "no secrets on disk" is only as good as the paths checked; and when a plaintext leak is found, expect a second, self-renewing source — anything the agent exported into its own shell, or any log that re-records the command line, reappears after redaction. Redaction closes what you found; only rotation closes what keeps regenerating, so say which one you achieved.
  • Locate the renewal mechanism by mtime, not by re-scanning content. Sort the suspect files by mtime: if the newest still contains the secret while older siblings you already redacted are clean, the source is live and re-emitting — a leftover is the opposite signature (old file, dirty). That single comparison converts "redaction failed" into "the emitter is a process I don't control", which is the difference between a fixable task and an operator action.
  • On a desktop agent gateway, that emitter outranks your shell. Gateway/terminal sessions snapshot the persistent shell's entire declare -x environment to a snapshot file on disk (e.g. cache/terminal/<host>-snap-*.sh), so any secret present in that environment is written to disk on every snapshot. unset in your own shell does not help — the next session inherits the gateway's environment again. The durable fix is upstream of your shell: never export a secret; read it into a plain variable and inject it per command (KEY=$(grep … .env), then env KEY="$KEY" cmd), so it never enters the process environment at all.
  • Your own re-runnable verification script is a leak vector, not just a test. A sweep script that exports a secret at the top runs in the persistent shell, so the export outlives the script and re-poisons every later snapshot — each "clean" verification run re-creates the leak it is checking for. Audit the harness for export SECRET with the same suspicion as the product. After changing it, re-run the script and confirm the variable is absent from the environment; that check is the proof the fix took.
  • Find the vendor's agent-integration documentation before reverse-engineering the integration. Vendors of decision/guardrail services often publish an agent-facing skill or doc index whose first rule is "read the API page before writing an integration", and individual pages may specify the exact component you are about to hand-debug. Enumerate the doc index (llms.txt or equivalent) as step one of any third-party integration; a cookbook matching your component turns a multi-round error-probing session into one page read, and tells you which of your components are faithful ports that must NOT be "fixed".
  • Close with a self-assessment that separates falsified-in-both-directions confidence from claims you cannot earn (provenance and correctness criteria belong to the independent reviewer, never the author), and disclose your own false-green or wrong-path detours — they are the reviewer's best guide to which tests still lack power.