Vesara
Vesara Daily Tuesday, September 15, 2026
 
What's in today's edition
Pion, an agent designed to run any company autonomously. Why do machine learning research agents not overfit?. On the funding desk: Zero (0.3m). Also Chift (€10.5m). Research watch covers the day's strongest operator-fit papers.
 
01  Agent abilities Skills · MCPs
 
Skills
01 Architecture decision records
Creates architecture decision records so a team can document a choice, its context, and the trade-offs behind it instead of leaving design rationale scattered through chats and pull requests.
Why it matters: Useful when agent-built systems need an audit trail for decisions that will outlive the current implementation.
Engineering practice High
02 Google Agents CLI workflow
Guides work through the Google Agents CLI, giving an agent a repeatable way to run the CLI rather than treating each task as a one-off command sequence.
Why it matters: Worth testing if your team wants a consistent harness around Google agent work instead of prompt-only operating habits.
Agent workflow High
03 Vitest
Adds a focused testing workflow for Vitest, helping an agent work with the JavaScript test runner while writing, running, and diagnosing tests.
Why it matters: A practical guardrail for code agents: tests catch the plausible-looking changes that miss the actual requirement.
Testing High
04 Astro
Provides an Astro-specific skill so an agent can work with the web framework using its conventions rather than applying generic React or static-site assumptions.
Why it matters: Useful for teams shipping content-heavy sites where framework-specific context saves cleanup work later.
Web development Medium
05 Security audit
Frames a security-audit task for an agent, giving it a defined way to inspect an application or change set for security issues before work ships.
Why it matters: Good fit for an approval step before automated code changes reach production systems.
Security High
 
MCPs
01 DC Hub
Gathers live signals spanning power networks, energy markets, and data-center operations so an agent can answer infrastructure questions with a common operational view.
Why it matters: Relevant for teams operating compute-heavy workloads where power and grid constraints are becoming part of planning.
Infrastructure intelligence Medium
02 Local Model Suitability MCP
Checks whether a task should use local Ollama inference or cloud inference before execution, acting as a routing layer for model choice.
Why it matters: Useful for controlling cost, latency, and data exposure without forcing every workflow onto one inference path.
Model routing High
 
02  Trending repos GitHub · last 24h
 
01 rlaope/oh-my-hermes  MIT — An all-in-one plugin for Hermes Agent that packages extra capabilities around the coding-agent environment. (+ today) 2,216 ★
02 666ghj/MiroFish  AGPL-3.0 — A swarm-intelligence engine aimed at modeling and predicting outcomes, putting multi-agent simulation into a project you can inspect and run. (+ today) 73,404 ★
03 tech-leads-club/agent-skills  No license — A registry for professional agent skills that focuses on secure, validated entries rather than an uncurated pile of prompts. (+ today) 6,163 ★
04 microsoft/AI-Engineering-Coach  MIT — A Microsoft project for better agentic engineering, aimed at improving how teams build software with AI agents in the loop. (+ today) 4,165 ★
05 NangoHQ/nango  No license — An integration platform for products using AI, giving builders a way to connect external services without rebuilding each connector. (+ today) 12,111 ★
Trending ≠ vetted. Star counts measure attention, not safety — review the code and pin versions before running anything marked early.
 
03  AI & tech news Key reads · Top 5
 
01 Pion, an agent designed to run any company autonomously  HN
Andon Labs describes Pion as an agent intended to run a company autonomously. It is a concrete glimpse at how far founders want to push delegation beyond a single task or department.
02 Why do machine learning research agents not overfit?  HN
Amazon Science asks why ML research agents avoid overfitting, a question that matters when an automated research loop starts optimizing against its own narrow measurements.
03 Notes on migrating 35KB preprompts from Opus to self-hosted Ollama  HN
A practitioner documents moving large preprompts away from hosted models to self-hosted Ollama. The useful part is the operational friction that appears when prompt context meets local serving.
04 An AEF-1 standard emerges for third-party evaluators  Latent Space
Latent Space reports that xAI, OpenAI, and Anthropic have all cosigned the AEF-1 standard for third-party evaluators, putting more structure around external AI assessment.
 
04  Reddit watch Top 5
 
01 15ms at P50 memory retrieval does absolutely nothing for a voice agent  R/AI_AGENTS
A practitioner challenges memory-tool marketing based on retrieval latency, arguing that a low P50 figure alone says little about whether a voice agent feels useful.
02 What AI agents are good for contact centers?  R/AI_AGENTS
A team operating high-volume support asks where agents genuinely help in contact centers, a grounded prompt to separate useful automation from vendor demos.
03 Fully offline RAG over 27 technical books  R/RAG
A builder shares an offline RAG setup across 27 technical books and 11,000 pages, with gold-set measurements of retrieval changes rather than just impressions.
04 The retrieval was fine, the model was quietly ignoring half of what I sent it  R/RAG
A RAG debugging post traces confident wrong answers to the model ignoring part of the supplied context, even though retrieval itself looked healthy.
05 GPT-5.6 Luna vs GPT-6 Astra on 50 real PRs  R/CHATGPTCODING
A developer publishes a comparison of GPT-5.6 Luna and GPT-6 Astra across 50 real pull requests and asks for scrutiny of the evaluation method.
 
05  Funding
 
01 Zero · 0.3m  Seed
Zero raised a 0.3m seed round backed by Lovable, Supercell, and Langdock founders, according to Sifted.
02 Chift · €10.5m  Series A
Brussels-based Chift raised a €10.5m Series A to build a European financial-connectivity layer.
 
06  Research watch Daily top
 
01 Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation  arXiv · September 15, 2026
A paper proposing a living database and search engine for AI benchmarks and evaluation, aimed at researchers and developers working with LLMs and other AI systems.
02 ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search  arXiv · September 15, 2026
Introduces ZGCM-1, an open 7B dense model trained for mathematical work and agentic search, making it the most directly operator-relevant model release in this set.
03 SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking  arXiv · September 15, 2026
Studies post-training attention sparsification to reduce the quadratic cumulative attention cost by optimizing which context receives attention.
 
  Your feedback
 
Reply to this email with your feedback, and we'll do our best to implement it in the next edition.
 
Vesara mark Vesara Find the 5% of revenue leaking out of your company.
InstagramLinkedInX Back to vesara.ai ->
Vesara Daily · Curated by Vesara operators · © 2026 Vesara, Inc. · vesara.ai Unsubscribe