|
|
|
|
What's in today's edition
Real-SWE benchmarks models on private enterprise code: Real-SWE evaluates AI models against private, real-world enterprise codebases. Dario Amodei argues for pacing frontier development: Anthropic CEO Dario Amodei lays out the case. On the funding desk: The Exploration Company (EUR 387M). Also Bending Spoons / Miro (EUR 1.7B).
|
|
|
| |
|
01 Agent abilities
|
Skills · MCPs |
Skills
| 01 |
google-agents-cli-deploy
Adds deployment steps to the Google Agents CLI, moving an agent project from local work into a deployed environment through the same command-line workflow.
Why it matters: Useful when an agent build needs a repeatable handoff from development to a running service instead of a separate deployment checklist.
Deployment
Delivery
|
| 02 |
reproduce-bug-report
Turns a reported defect into a reproducible case by working through the product flow, giving an agent a concrete way to investigate vague bug reports.
Why it matters: A reproducible case narrows the work before someone starts changing code and makes a fix easier to verify.
Engineering
Debugging
|
| 03 |
sentry-cli
Provides a command-line path into Sentry so an agent can work with error-monitoring tasks without leaving the terminal workflow.
Why it matters: It bridges an incident signal and the commands an engineering agent can run while investigating it.
Observability
Reliability
|
| 04 |
sense
Offers a skill for sensemaking, helping an agent organize and reason through a body of information before it produces an answer or plan.
Why it matters: Worth testing for research and operating reviews where the hard part is sorting evidence, not generating more text.
Research
Judgment
|
| 05 |
agent-harness-construction
Guides the design of an agent harness, including its action space, tool definitions, and the observations supplied back to the model.
Why it matters: The harness determines what an agent can see and do. Tightening it often matters more than swapping the underlying model.
Agents
Architecture
|
MCPs
| 01 |
Playwright
Connects an agent to browser automation through Playwright, covering real-page interaction and browser-driven checks instead of relying on static HTTP requests.
Why it matters: Useful for workflows that must verify what a customer or operator actually sees in a browser, including interactive flows.
Browser
Automation
+2,024 pulls on Docker Hub since yesterday
|
| 02 |
SonarQube
Exposes SonarQube to an agent so code-quality findings can be inspected alongside the repository work that caused them.
Why it matters: It gives coding agents a route to catch quality and maintainability issues before a change reaches review.
Code quality
Engineering
+1,361 pulls on Docker Hub since yesterday
|
|
|
| |
|
02 Trending repos
|
GitHub · last 24h |
| — | No new trending repos today. |
Trending ≠ vetted. Star counts measure attention, not safety — review the code and pin versions before running anything marked early.
|
|
| |
|
03 AI & tech news
|
Key reads · Top 5 |
| 04 |
Simon Willison on using OpenRouter
Simon Willison
Willison examines OpenRouter and its pitch that it handles access across model providers for developers who do not want to manage each one separately.
|
|
|
| |
|
| 01 |
Kimi routed to Claude
R/CLAUDEAI
A user flags an experience where Kimi traffic was routed to Claude, prompting discussion about model routing and what users can infer from it.
|
| 05 |
How to claim webhook work safely
R/N8N
An automation-backend discussion covers duplicate webhook side effects and stuck jobs, then asks how teams safely claim work before processing it.
|
|
|
| |
|
| 01 |
The Exploration Company
· EUR 387M
Funding
The European space company follows its EUR 387 million raise with an ESA cargo-capsule contract reported at EUR 760 million.
|
| 02 |
Bending Spoons / Miro
· EUR 1.7B
Acquisition
Bending Spoons entered a definitive agreement to acquire the AI innovation workspace Miro for EUR 1.7 billion.
|
| 03 |
Mecka AI
· ~$500M valuation
Deal
Robot-training-data startup Mecka AI is reportedly nearing a $500 million valuation in a Sequoia-led transaction.
|
|
|
| |
|
06 Research watch
|
Daily top |
|
|
| |
|
Reply to this email with your feedback, and we'll do our best to implement it in the next edition.
|
|
| |
|
|