Evidence snapshot · 2026-08-18
Selected source notes
These notes spotlight the sources most useful for evaluating the paper’s empirical, protocol, security, and cost claims. The canonical white paper carries the complete numbered bibliography and places population and method limits beside each material claim.
Citation numbers retain the identities used in the white paper. A missing number here means only that this companion is selective; it does not create a new source identity.
Architecture and evidence foundations
- [1] Cerf, V., and Kahn, R. “A Protocol for Packet Network Intercommunication.” The historical foundation for the paper’s narrow-waist systems framing.
- [13] Kubernetes documentation: “Controllers.” Defines reconciliation as a control-loop pattern for declared and current state.
- [21] NIST AI Risk Management Framework and Generative AI Profile. Provides risk-management context; it does not certify Conexus.
- [22] NIST Secure Software Development Framework. Supports the secure-development and human-oversight framing.
Agentic software evidence
- [55] Mazloomzadeh, Morovati, and Khomh. Agentic pull-request binomial GLMM over 9,227 pull requests. The model reports an estimated merge-probability range; it is not a controlled human-versus-agent benchmark.
- [57] Ehsani et al. “Where Do AI Coding Agents Fail?” A 33,596-pull-request snapshot used with its published abandonment denominators.
- [58] “Why Are Agentic Pull Requests Merged or Rejected?” A manual analysis of 353 rejected pull requests; conclusions are scoped to that sample.
- [59] Duma et al. Agentic pull-request review populations. Separates review-population evidence from broader acceptance claims.
- [60] DORA. “Balancing the Tensions of AI-Assisted Software Development.” Organization-level evidence about AI adoption, delivery, and system conditions.
- [61] Stack Overflow Developer Survey 2025, AI section. The paper reports response counts beside the distinct trust and accuracy questions.
- [63] Zhong, Raghunathan, and Carlini. “ImpossibleBench.” The paper distinguishes scaffold, read-only, and human-escalation experiments.
- [65] Tang et al. Developer-agent misalignment in 20,574 real-world sessions. The paper keeps sessions, episodes, visible resolutions, and after-pushback behavior as separate denominators.
- [67] METR. “Frontier Risk Report.” Time-horizon estimates are presented with their intervals and task boundaries.
Context and cost evidence
- [70] Anthropic Claude Code worked cost example and model pricing. The paper labels derived arithmetic as a derivation rather than a quoted provider statistic.
- [71] Bai et al. Eight-model OpenHands and SWE-bench Verified token-economics study. Supports benchmark-specific token-use analysis, not a universal cost multiplier.
- [73] Rafiei Oskooei et al. Deep agentic search for repository-level code question answering. Evidence for measured repository retrieval within the study’s task design.
- [87] Anthropic. Building effective agents with managed context. Supports the persisted-context boundary described in the paper.
- [92] Anthropic model pricing. A dated provider source for long-context rates; rates can change.
Protocols, provenance, and security
- [77] Model Context Protocol HTTP Authorization, revision 2026-07-28. Authorization exists for HTTP transports and remains optional at the protocol level.
- [78] Agent2Agent Protocol 1.0, Agent Card signing. Signing can establish card integrity and provider origin when verified against a trusted key; it is not universal agent identity.
- [79] MCP OAuth-enabled server measurement. A May 2026 measurement study whose results remain bounded to the sampled ecosystem.
- [80] NIST SP 800-218 and SP 800-218A. Secure development guidance used as design evidence, not as a compliance assertion.
- [81] SLSA Build Provenance specification v1.2.
- [82] in-toto Attestation Framework statement specification.
- [83] Ollama Cloud documentation. Supports the paper’s description of cloud execution through Ollama interfaces.
- [84] Tan, Garry. gstack, an MIT-licensed multi-agent workflow suite.
- [91] Model Context Protocol security best practices. A current draft, identified as such in the paper.