|
|
|
|
What's in today's edition
OpenAI Agents API. Cognition launches SWE-2. On the funding desk: Seed Capital (€130M). Also Furo (€3.44M). Research watch covers the day's strongest operator-fit papers.
|
| |
|
01 Agent abilities
|
Skills · MCPs |
Skills
| 01 |
p-image
An image-handling skill for agent workflows. It looks like a small building block for teams that need repeatable visual outputs.
Why it matters: If your agents produce creative assets, a clear handoff between text and image work keeps the process reviewable.
skills.sh
Creative ops
Content
|
| 02 |
Decision records
A discipline for writing down technical and product choices as they are made, rather than retrying to reconstruct them from chat logs.
Why it matters: A decision log makes a client agent easier to operate, debug, and hand over when the team changes.
skills.sh
Engineering ops
Reliability
|
| 03 |
Go project layout
A focused guide for structuring Go projects when an agent is creating or changing services. It makes architectural conventions available at the point where code gets written.
Why it matters: Useful for keeping an agent-generated service legible to the next engineer instead of accumulating a custom layout one task at a time.
skills.sh
Engineering
Code quality
|
| 04 |
Svelte 5 best practices
A guide for agents working in Svelte 5 codebases. It puts framework practices next to the work so common missteps are less likely.
Why it matters: A targeted skill can prevent a coding agent from retrieving generic advice and leaving a framework-specific mess behind.
skills.sh
Front end
Build quality
|
| 05 |
Video ad specs
A practical reference for turning a campaign brief into video-ad requirements an agent can follow. It narrows the gap between a creative request and assets that can actually ship.
Why it matters: Useful when campaign work reaches production and the agent needs guardrails around delivery requirements, not only ideas for the creative.
skills.sh
Creative ops
Marketing
|
MCPs
| 01 |
URL Safety Validator
A gate that checks links for phishing and malware before an agent opens them. It fits unattended browsing flows.
Why it matters: Add it before a browser agent follows lead-research links or opens untrusted attachment routes.
Smithery
Security
Risk control
|
| 02 |
Niche
An editorial-research tool for finding sourced angles and ranking topics before the content team writes.
Why it matters: It should help turn the daily from a list of links into a few ideas worth actually writing or testing.
Smithery
Content ops
Editorial
|
| 03 |
ReportFlow
A connector that turns existing report templates into PDFs such as invoices, contracts, and statements from an MCP-compatible assistant. It focuses on producing a finished document from a controlled template.
Why it matters: Useful when client ops repeatedly turn structured work into PDFs and you want the agent to preserve the approved layout rather than invent one.
Smithery
Document ops
Client ops
|
| 04 |
OrgX
A coordination layer for AI-native teams with initiatives, milestones, tasks, delegated work, and persistent organizational context. It is aimed at keeping execution connected to an operating plan.
Why it matters: Worth testing where client agents cross sessions and owners need a readable view of decisions, work in flight, and approval points.
Smithery
Coordination
Operator ops
|
|
| |
|
02 Trending repos
|
GitHub · last 24h |
| 01 |
Tencent/teamai-cli
No license
— A TCL for making teams AI-native. It is a fresh open-source option for teams trying to coordinate agent work across an organization. (+841 today)
|
4078 ★ |
| 02 |
JustVugg/colibri
Apache-2.0
— A pure-C engine for running mixture-of-experts models on existing hardware by streaming experts from disk. It is one more sign that local inference constraints are moving quickly. (+98 today)
|
27625 ★ |
Trending ≠ vetted. Star counts measure attention, not safety — review the code and pin versions before running anything marked early.
|
| |
|
03 AI & tech news
|
Key reads · Top 5 |
| 01 |
OpenAI Agents API
HN
OpenAI's new Agents API brings agent building into its developer docs. Read the surface area carefully: the operational question is how tools, state, and evals behave under real workloads.
|
| 02 |
Cognition launches SWE-2
HN
Cognition introduced SWE-2 as a new coding model. The claim is worth testing against your own bug-fix, feature, and regression workloads rather than benchmarks alone.
|
| 03 |
Anthropic reports on AI misuse
HN
Anthropic describes campaigns that abused models for attacks. For anything that browses or takes action, the operational lesson is to keep permissions, logs, and escalation paths explicit.
|
| 04 |
Evals for everyone
Every
Every makes the case for evaluations that a regular team can actually use, not just research groups. The timing is right: agent quality without a test set is largely a feeling.
|
|
| |
|
| 01 |
Claude to reMarkable
R/CLAUDEAI
A user shows Claude turning to-dos and follow-ups into a daily worksheet for a reMarkable device. It is a tiny, concrete example of an agent fitting into a human workday.
|
| 02 |
Harness does matter
R/LOCALLLAMA
A practitioner redeciscovers that the setup around a flash model can change the result more than expected. The useful lesson is to compare harnesses, not just model names.
|
| 03 |
OUI-1 generates bespoke UI
R/LOCALLLAMA
A post examines OUI-1, a model trained to make custom interface elements via a specialized DSL. It offers a different path for prototyping uis without starting with raw React code.
|
| 04 |
Codex vs Claude on a creative UI brief
R/CLAUDECODE
The same open-ended game brief was handed to two coding assistants. It is a reminder to examine the workflow and build quality, not just the first screenshot.
|
|
| |
|
| 01 |
Seed Capital
· €130M
Fund V
Copenhagen's Seed Capital closed a €130 million fund for Nordic seed to Series A investments in fintech, cybersecurity, resilience, and AI-led B2B.
|
| 02 |
Furo
· €3.44M
Venture round
Munich-based Furo raised €3.44 million to expand its industrial battery-storage software and enter more European markets.
|
| 03 |
CloudNC
· $20M
Series B extension
CloudNC secured a $20 million Series B extension for manufacturing automation, taking total funding to $128 million.
|
| 04 |
Bluecore Energy
· $50M
Seed
Bluecore Energy raised a $50 million seed round for a nuclear energy startup, following a $10 million pre-seed and a stealth launch two months ago.
|
|
| |
|
06 Research watch
|
Daily top |
| 02 |
Miles v0.1: Production-Level Post-Training
HF Papers · 50 upvotes · Sep 8
Miles describes a verifiable reinforcement-learning post-training stack, with components intended to be clean and customizable. It is a useful infrastructure reference if you're building evaluations or finetunes.ing.
|
|
| |
|
Reply to this email with your feedback, and we'll do our best to implement it in the next edition.
|
| |
|
|
|