Skip to content
Trending

Worth knowing

Ideas that outlast the news cycle. There are no dates on this page: each idea was true when we checked it and will still be true next quarter. Not everything that spreads on social media is true, so every claim is checked against its primary source and explained in our own words.

17 ideas · explained, then checked

Each card says what is going around, what is actually true, the one line to remember and what it means for a product. Open the checks to see the source behind every claim.

73 claims · 30 hold · 34 partly · 8 wrong · 1 opinion

Tools

Agent framework tier lists are opinion: of 4 dismissed projects, only AutoGen's README steers users away.

Tier lists of agent frameworks for production are going around. They put a single framework in the top tier, call LangGraph the best fit for controlled, sequential work in regulated industries but weak for fully autonomous tasks, dismiss CrewAI, AutoGen, Haystack and Pydantic AI, and rate building from scratch highest for control.

  • 3 layersruntime, framework, harness: how LangChain's docs sort what tier lists rank together
  • 1 of 4dismissed projects whose own README steers users elsewhere (AutoGen)
  • 3primitives in the OpenAI Agents SDK
  • 7durable execution engines Pydantic AI lists
  1. Agent frameworks can be ranked in one list, with a single framework in the top tier for production.A tier is a judgement, and none of the project pages cited here publishes a head-to-head measurement on a shared production task. The projects do not even claim the same job. LangChain's documentation sorts the field into 3 layers: runtimes that keep an agent running (LangGraph, beside Temporal and Inngest), frameworks that supply abstractions and integrations (LangChain, CrewAI, the OpenAI Agents SDK, Google ADK and LlamaIndex among them) and harnesses that ship tools, prompts and subagents (Deep Agents and the Claude Agent SDK among them). That sorting is one vendor's view, and the projects describe themselves differently again. Agno calls itself a framework and runtime for agent platforms, with a REST API, storage in your own database and a control plane. The OpenAI Agents SDK calls itself a lightweight package with 3 primitives: agents, handoffs (or agents used as tools) and guardrails. Pydantic AI calls itself a typed agent loop. CrewAI, Haystack, Pydantic AI, Microsoft Agent Framework and the OpenAI Agents SDK all describe themselves as production-ready or production-grade, so the label separates nobody.Runtimes, frameworks, and harnesses (LangChain docs)
  2. LangGraph is the best fit for controlled, sequential work in regulated industries.The control half matches what LangGraph says about itself: a low-level orchestration runtime for long-running, stateful agents, where hand-coded deterministic steps and model-driven steps sit in one graph, so some of the logic stays predictable and open to audit while the rest is left to the model. It lists persistence, so a run that fails can pick up where it stopped, and human-in-the-loop review, where a person can read and edit the agent's state mid-run. Two details do not come from the documentation. Sequential undersells it: LangGraph's workflows guide builds parallel calls, routing, orchestrator-worker and looping tool-calling agents as graphs. And the overview never mentions regulated industries; it names Klarna, Uber and J.P. Morgan as users, which is the vendor's own list and says nothing about compliance. Best is the ranking's opinion.LangGraph overview
  3. LangGraph is weak for fully autonomous tasks.The true part: LangGraph describes itself as a runtime and leaves the agent to you. Its overview says it does not abstract prompts or architecture, and it sends people who want a higher-level start to LangChain's prebuilt agents, so an autonomous agent on bare LangGraph is more code to write. That is not a ceiling on autonomy. LangChain's agents are built on top of LangGraph. LangChain's product comparison points teams building more autonomous agents for complex, non-deterministic tasks to a harness, and its own harness, Deep Agents, adds a virtual file system, subagents and an optional to-do list for planning. The Deep Agents overview says it uses the LangGraph runtime for durable execution, streaming and human review. The autonomy comes from the layer above; the runtime underneath is the same one.Deep Agents overview (LangChain docs)
  4. CrewAI, AutoGen, Haystack and Pydantic AI are not worth considering for production.1 of the 4 dismissals has a primary source behind it. AutoGen's own README says the project is in maintenance mode, will receive no further features, is community managed, and that anyone starting out should use Microsoft Agent Framework instead. Microsoft describes that framework as the direct successor to both AutoGen and Semantic Kernel, built by the same teams, with graph-based workflows, checkpointing and human-in-the-loop support. The other 3 carry no such notice in their READMEs or documentation, and each describes a specific job. CrewAI tells production users to start with a Flow, a structured, event-driven workflow that holds state, and to call a Crew of autonomous agents only inside a step that needs one. Pydantic AI lists typed outputs and tools, durable execution on 7 engines including Temporal and DBOS, and built-in human approval. Haystack, from deepset, builds retrieval and agent applications as pipelines of reusable components. Whether any of them suits your system is a test to run, and no tier list has run it for you.AutoGen README (maintenance mode notice)
  5. Building from scratch ranks highest because it gives the most control.Anthropic's guidance backs the starting point. It advises calling the model API directly first, says many agent patterns take a few lines of code, and warns that frameworks add layers which can hide the underlying prompts and responses, make debugging harder and tempt teams into complexity they do not need. It names wrong assumptions about what a framework does under the hood as a common source of customer error. The same article is not against frameworks: it lists 4, Anthropic's own Claude Agent SDK among them, says they make it easy to get started, tells teams that use one to understand the code underneath, and encourages cutting abstraction layers on the way to production. Control also has a bill. By their own documentation, saved state that survives a restart and a pause for human approval come built into LangGraph, Microsoft Agent Framework and Pydantic AI on a durable engine; a from-scratch build has to write and maintain each one it needs, plus its own tracing. OpenAI and Anthropic both present it as a choice of who owns the loop. OpenAI's SDK docs say to call the Responses API directly when you want to own the loop, tool dispatch and state, to use the SDK when you want the runtime to manage them, and that many applications do both. Anthropic's SDK docs draw the same line between its Agent SDK, which runs the agent loop for you, and its Client SDK, where you write the tool loop yourself.Building effective agents (When and how to use frameworks)

the line to remember

Pick the layer before the brand: start with direct model calls, add a runtime, framework or harness only for what it ships that you need, and read any tier list as opinion.

For your product

Start from what the system must survive. A process with fixed steps, sign-off points and an audit trail needs explicit workflow control, saved state and human approval, which LangGraph, CrewAI Flows, Microsoft Agent Framework workflows and Pydantic AI on a durable engine all document. An open-ended task needs an agent loop with tools, and most of the same projects offer one as well. Ask a vendor 3 things: which parts of the flow are fixed in code and which are left to the model, what happens to a run when the server restarts halfway, and whether your own team can read and debug the prompts the framework sends. Check the maintenance status in the project's own repository before you commit; AutoGen's README shows why. A prototype on direct API calls is cheap and tells you which of these features you need before you adopt anything. Microsoft's own guide adds the cheapest option of all: if an ordinary function can handle the task, write the function and skip the agent.

Tools

Agent Beacon logs 8 of 8 signal types for Claude Code, 6 of 8 for Codex CLI, and blocks nothing by default.

A free open-source tool called Agent Beacon is being passed around as the way to see what an AI coding agent does on your laptop. The description: a background service and a dashboard on a localhost port that list every command, tool call and file edit from Claude Code, Cursor and Codex as they happen, plus rules that raise an alert when an agent reads something sensitive.

  • 29local agent runtimes in Beacon's coverage table
  • 6 of 8signal types Beacon marks for Codex CLI, against 8 of 8 for Claude Code
  • 6Claude Code permission modes; deny rules block in all of them
  • 512 charscap per value in Claude Code's exported tool input

the line to remember

A log of what your coding agent reports is worth keeping, but it is the agent's own account written after the fact, so put secrets behind deny rules and a sandbox first and use the log to check that they held.

For your product

Treat this as two decisions. Visibility: Claude Code and Codex can export the trail over OpenTelemetry themselves, and Claude Code's telemetry settings can be pushed through managed settings, so a team with a collector and a SIEM may need no extra software on each laptop. A tool such as Agent Beacon earns its place when developers use several agents and you want one format, a local viewer and ready-made rules. Decide what may be stored before enabling content, because commands and file paths can carry secrets. Anthropic keeps prompts and tool inputs out of the export unless you switch them on, OpenAI redacts prompt text by default and tells teams to treat tool arguments and outputs as sensitive, and Beacon's Claude Code page says its install turns detailed tool logging on. Read Beacon's installer prompts as well: the README presents the project first as a memory layer for coding agents, which the description going around leaves out, and interactive setup signs in through beacon.sh with the hosted option preselected, though the README states that nothing is forwarded until a separate connect step and that package, MDM and CI installs need no account. Prevention: require the sandbox, switch off the unsandboxed fallback in Claude Code, deny reads of credential folders, set security hooks to fail closed, and deliver those settings through managed or team policy so neither a developer nor a repository file can loosen them.

Roadmap

13 papers behind AI engineering, 5 one-line summaries checked: 1 holds, 2 are half right, 2 are wrong.

A short reading list of the papers behind day-to-day AI engineering is widely shared: Attention Is All You Need, BERT, ViT, GANs, VAEs, diffusion models (DDPM), LoRA, RAG, mixture of experts (Switch Transformer and Mixtral), RLHF and InstructGPT, LLaMA and RoPE. Each title tends to travel with a one-line summary, for example that LoRA trains 0.01 percent of the weights, or that a 1.3B model beat a 175B one.

  • 0.01%of GPT-3's weights trained by LoRA at rank 4, 18M of 175B
  • 135×size gap the 1.3B InstructGPT overcame on human preference, not on benchmarks
  • 47B / 13BMixtral parameters in memory and per token, not 56B and 7B
  • 3.5 daysto train the big Transformer on 8 P100 GPUs

the line to remember

The reading list is sound, but the one-line summaries keep the headline number and drop its condition: LoRA cut training memory 3 times, not 10,000; the 1.3B model won on preference, not on benchmarks; Mixtral needs memory for 47B parameters, not 13B.

For your product

Several of these papers sit directly behind line items in an AI budget, so the misread versions cost money. LoRA cuts what you store per customised model (35MB against 350GB in the paper's GPT-3 example), yet the GPUs must hold the full base model, so it does not turn a large model into a small one. A mixture-of-experts model costs compute by its active parameters and memory by its total parameters: size the hardware for Mixtral's 47B, not its 13B. A small tuned model can beat a far larger one on what users prefer, which is a reason to test on your own prompts with your own raters, and to re-run capability tests after tuning, because the InstructGPT authors measured regressions on 4 named public benchmarks. When a vendor quotes one of these papers, ask which table the number comes from and what was held fixed.

Sources: Attention Is All You Need (Vaswani et al., Google Brain, Google Research and University of Toronto) · BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding (Devlin et al., Google AI Language) · An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (ViT) (Dosovitskiy et al., Google Research, Brain Team, ICLR) · Generative Adversarial Networks (Goodfellow et al., Université de Montréal) · Auto-Encoding Variational Bayes (VAE) (Kingma and Welling, Universiteit van Amsterdam) · Denoising Diffusion Probabilistic Models (DDPM) (Ho, Jain and Abbeel, UC Berkeley) · LoRA: Low-Rank Adaptation of Large Language Models (Hu et al., Microsoft) · Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., Facebook AI Research, UCL and NYU) · Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity (Fedus, Zoph and Shazeer, Google, Journal of Machine Learning Research) · Mixtral of Experts (Jiang et al., Mistral AI) · Training language models to follow instructions with human feedback (InstructGPT) (Ouyang et al., OpenAI) · LLaMA: Open and Efficient Foundation Language Models (Touvron et al., Meta AI) · RoFormer: Enhanced Transformer with Rotary Position Embedding (RoPE) (Su et al., Zhuiyi Technology) · Neural Machine Translation by Jointly Learning to Align and Translate (Bahdanau, Cho and Bengio, Jacobs University Bremen and Université de Montréal)

Concept

9 of 10 widely used benchmarks score 'I do not know' as zero, so language models learn to bluff.

The usual explanation is that a language model invents facts because it only predicts the next word, and that connecting it to your own documents makes the problem go away. A fuller version says training and testing reward a confident guess over an honest 'I do not know', and that the fixes with evidence are grounded answers with citations, permission to abstain, and scoring that punishes confident errors.

  • 9 of 10benchmarks with no credit for abstaining
  • 26% vs 75%wrong answers on SimpleQA, declining vs guessing model
  • 74.5%citations that support their sentence

the line to remember

Models bluff because bluffing scores well; score a wrong answer below 'I do not know', ground each claim in a passage, and check the passage supports it.

For your product

Accuracy alone is the wrong number to buy on. Ask a vendor for three figures measured on your own questions: how often the system is right, how often it is wrong, and how often it declines. In acceptance tests, score a wrong answer as worse than a declined one, include questions your documents cannot answer so you see what happens when search comes back thin, and sample the cited passages to confirm they support the sentence. A system that sometimes declines and is rarely wrong is usually worth more than one that always answers.

Sources: Why Language Models Hallucinate (Kalai et al., OpenAI and Georgia Tech) · Sufficient Context: A New Lens on Retrieval Augmented Generation Systems (Joren et al., UC San Diego, Duke University and Google) · Evaluating Verifiability in Generative Search Engines (Liu, Zhang and Liang, Stanford University, Findings of EMNLP) · Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks (Lewis et al., Facebook AI Research, UCL and NYU, NeurIPS) · GPT-5 System Card (OpenAI)

Concept

5 times the precision at near-equal recall in one open test: 200-token chunks against an 800-token default.

Guides to retrieval-augmented generation tend to hand out a recipe for splitting documents: choose a chunk size such as 512 or 800 tokens, add 10 to 20 percent overlap, and split on structure or meaning instead of fixed lengths. Contextual retrieval and late chunking are then presented as the fix for chunks that lose their surrounding context.

  • 5×precision, recursive 200-token chunks against the 800/400 default (Chroma)
  • 5.7% → 1.9%missed top-20 retrievals with contextual retrieval, BM25 and a reranker (Anthropic)
  • +1.8nDCG at 10 points from late chunking, 52.2 to 54.0

the line to remember

Chunking is a setting you measure on your own documents and questions, because the best splitter, size and overlap changed with the data or the embedding model in each study cited here.

For your product

Ask a vendor which chunk sizes and splitters they tested on your documents, and for the recall and precision of each. If the answer is that they use the default, nobody measured. Start with the cheap option, a paragraph-aware splitter at a few hundred tokens, build a test set of real questions paired with the passages that answer them, and pay for per-chunk context or embedding-based chunkers only when that test shows the gain.

Concept

Flat chunks throw away the table of contents. Keeping it was worth 5.7 points in IBM's STAIR test, not 23.1.

A retrieval method credited to IBM researchers is going around: instead of cutting documents into equal chunks and embedding each one, let the model use the document's own structure, its table of contents and section hierarchy, to find the right section. It is usually presented as a replacement for chunk-and-embed retrieval.

  • 5.7 ptsSTAIR over the same tuned model with no contents page
  • 13.9 ptsSTAIR over an untuned dense retriever, right section ranked first
  • 2.0 ptsRAPTOR's tree over a dense retriever, same reader, QuALITY

the line to remember

Headings are retrieval signal the author already wrote, so keep them through parsing and let search use them alongside embeddings, not instead of them.

For your product

Before paying for a bigger embedding model, check what your ingestion does to headings. Manuals, policies, contracts and textbooks arrive with a contents page; if the parser flattens it, every chunk loses its address. The cheap version needs no training: chunk along the document's own elements and store the heading path with each chunk, which the chunkers in Docling, an open-source parser started at IBM Research, do for you. None of the papers above measures that cheap version, so test it on your own questions. The trained versions (STAIR, RDR2) cost a fine-tuning run each, and when the comparison is fair the published gains are single digits: 5.7, 3.4 and 2.0 points, not 20. None of this helps a corpus with no structure, such as chat logs or email. Ask any vendor two things: does the pipeline keep section headings, and was the benchmark run on documents like yours with baselines tuned the same way.

Concept

A router decides once and steps aside. A supervisor keeps deciding until the job is done.

Multi-agent diagrams often draw a router and a supervisor as the same box: something at the top that picks which agent gets the work. The usual claim is that the two words are interchangeable because both choose an agent.

  • 4 vs 3model calls for one simple request, supervisor vs router
  • 90.2%supervised team over a single agent, Anthropic internal eval
  • 15×tokens of a multi-agent run compared with a chat

the line to remember

A router chooses once and leaves; a supervisor chooses, reads the result and chooses again until the work is done.

For your product

When a vendor draws one box above several agents, ask whether that box decides once or stays in charge. If requests fall into clear categories that one specialist can finish, a router is cheaper, faster and easier to test, because you can score it like any classifier. If a request needs several specialists and the next step depends on what the last one found, you need a supervisor, and a budget for its extra calls and for tracing its decisions. Amazon Bedrock makes the choice a single setting, SUPERVISOR or SUPERVISOR_ROUTER, allows up to 10 collaborator agents per supervisor and notes that the routing mode reduces latency.

Concept

The OAuth standard lists 5 problems with giving software your password. All 5 apply to AI agents.

Common advice for connecting an AI agent or an MCP server to email and other accounts: never give it the real account password. Use OAuth with narrow scopes, fall back to an app password only where OAuth is missing, keep every credential revocable, and for remote MCP servers rely on the authorisation model in the MCP specification.

  • 5problems with password sharing, RFC 6749
  • 14Gmail API scopes to choose from
  • 8 of 14Gmail scopes Google marks restricted

the line to remember

An agent should hold a token that names what it may do, expires, and can be switched off on its own; a password does none of the three.

For your product

Before an agent touches a mailbox, a drive or a CRM, ask the vendor three things: which scopes it requests and why each is needed, where the tokens are stored, and how you revoke one agent without locking out your staff. Treat any connect screen that asks for the account password in the product's own form as a failed review. For MCP servers, ask whether the server follows the specification's authorisation rules or reads a long-lived secret from a config file, because the protocol allows both.

Concept

About 100 tokens per installed skill: the rest of the folder stays on disk until a task matches.

Agent skills are described as a way to teach an agent a job once: a folder with a SKILL.md file holding instructions, scripts and resources, which the agent reads only when a task needs it. Common add-ons to that story are that a skill is just a saved prompt, that skills make tools and MCP servers unnecessary, and that a skill is only text, so it can be installed from anywhere.

  • ~100 tokensper installed skill, at rest
  • <5k tokensSKILL.md body, loaded on match
  • 3 levelsmetadata, instructions, resources

the line to remember

A skill is a folder of know-how that costs about 100 tokens until a task matches it, so write it as you would brief a colleague joining the team and vet it as you would any program you install.

For your product

The know-how that makes an agent useful in your company (how a report is laid out, which checks a refund needs, how a release is cut) can live in version-controlled folders that your team reviews like code. Because the format is an open standard, the same folder can be read by agent products from several vendors, which lowers the cost of switching, though bundled scripts still depend on what each environment allows (network access, installed packages). Put skills under the same controls as software: a named owner, review before install, and no unvetted downloads on machines that hold customer data.

Concept

33 percent less error than seasonal naive on 97 public tasks it never trained on: pretrained forecasters.

Forecasting is said to have its own foundation models: one pretrained model that predicts sales, traffic or sensor readings it has never seen, with no training on that data, and matches models built for each dataset. Google Research's TimesFM is the usual example, described as open weights, able to read about 16,000 past points, and first on every public benchmark.

  • 33%lower point error than seasonal naive, zero-shot (GIFT-Eval)
  • 330Mparameters, version 3.0 (200M in 2.5)
  • 15,360past points read at most
  • 1T+time points in pretraining, up from 100B in the first model

the line to remember

A pretrained forecaster is a strong baseline you get before building anything, so backtest it zero-shot on your own history beside seasonal naive, and make any custom model beat both before you pay for it.

For your product

A first forecast no longer needs a data science project per product line. The same model family runs inside BigQuery as AI.FORECAST with no model to train. Check three things before relying on it: the licence (3.0 weights are non-commercial, 2.5 is Apache-2.0, a managed cloud service has its own terms), a backtest on your own history against seasonal naive and your current method, and whether your real drivers such as price, promotions and holidays can be passed in as covariates. If the backtest will inform a business decision, run it on the Apache-2.0 version or the managed service, because the non-commercial licence rules that use out. Public benchmarks rank models on public data; your data decides.

Concept

Elastic weight consolidation puts a spring on each weight, stiffest where an old task would break.

Elastic weight consolidation (EWC) is usually presented as the cure for catastrophic forgetting, the way a neural network loses an old task while it is trained on the next one. The common account: after the first task, score every weight by its Fisher information, then penalise changes to the high-scoring weights during later training. It is often said to let one network keep learning task after task with nothing lost.

  • 3numbers stored per weight under EWC: value, anchor, importance
  • 20.01%EWC on split MNIST when it must tell tasks apart
  • 90.79%generative replay on the same test

the line to remember

EWC charges the network for moving the weights an old task depends on, which slows forgetting without ending it, and replay still beats it once the network has to tell the tasks apart by itself.

For your product

When a supplier fine-tunes a model on your data, ask what it lost as well as what it gained. A study of language models from 1 to 7 billion parameters found forgetting was the general pattern when instruction tasks were tuned in one after another, and it grew worse with model size in that range, so insist on before-and-after scores for the general skills you depend on. The usual protections are a penalty such as EWC (two extra numbers held per weight, plus a search for the right strength), mixing some earlier data back in, or training a small adapter such as LoRA, which one comparison on code and maths found forgets less than full fine-tuning and also learns less.

Sources: Overcoming catastrophic forgetting in neural networks (Kirkpatrick et al., DeepMind and Imperial College London (PNAS)) · On Quadratic Penalties in Elastic Weight Consolidation (Ferenc Huszár) · Three scenarios for continual learning (van de Ven and Tolias, Baylor College of Medicine) · Full-Parameter Continual Pretraining of Gemma2: Insights into Fluency and Domain Knowledge (Šliogeris et al., Neurotechnology) · An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning (Luo et al., Tencent WeChat AI and Westlake University) · LoRA Learns Less and Forgets Less (Biderman et al., Columbia University and Databricks Mosaic Research)

Interview question

10 users can crash a 48 GB GPU running a 13B model. The maths says 60 GB.

An interview-style post: a 13-billion-parameter model is deployed on a 48 GB GPU, only 10 users are on it, and the server still dies with CUDA out of memory. The caption blames the KV cache, uncontrolled concurrency and memory fragmentation, and prescribes continuous batching (vLLM, TensorRT-LLM, TGI), capping tokens and context, keeping 10 to 15 percent of memory free, and queueing with backpressure.

  • 26 GB13B weights, fp16
  • 0.8 MBKV per token, 13B
  • 60 GB10 users × 4k tokens

the line to remember

Serving memory = weights + (tokens × concurrent users × per-token KV cost). Size the card for the cache, not the model.

For your product

If a vendor quotes you a GPU from the model size alone, the quote is wrong. Ask for the per-token KV cost, the context cap and the concurrency cap; those three numbers decide whether the box is enough.

Interview question

All weights set to 0: every neuron gets the same gradient, so the network never learns.

A whiteboard post asks whether you can initialise all weights to zero. Answer: no, the gradient is identical for every neuron, so weights must be random to break symmetry. A commenter adds that with no bias term everything stays exactly zero; another says use Glorot for sigmoid or tanh and He for ReLU.

the line to remember

Initialisation is about breaking symmetry and keeping signal variance stable through the layers. Zero does neither.

For your product

Nothing to decide here unless you train models. If you do, this is the first question a reviewer will ask about any training bug that ends with a network that outputs the same answer for everything.

Concept

3 agent patterns on one graphic: CodeAct, ReAct and agentic RAG. Two of the three descriptions oversell.

A graphic contrasts a single agent with a multi-agent system, then defines CodeAct (the agent acts by writing and running Python), ReAct (reasoning traces interleaved with tool actions, said to overcome hallucination and error propagation) and agentic RAG (agents orchestrating the retrieval pipeline).

the line to remember

ReAct is the loop, CodeAct is the action language, agentic RAG is the loop pointed at your documents. Pick the simplest one that fits the task.

For your product

Most business tasks do not need a multi-agent system. A single agent with a few well-described tools and a stop condition covers a lot; add agents only when the path cannot be known in advance.

Research

91.2 percent: how often innocent tool calls could be chained into a harmful action in Amazon's STAC study.

Two papers led by Amazon interns were accepted: one on relational priors in LLM multi-agent systems (AACL) and STAC, on how benign tools can form dangerous chains for LLM agents (EMNLP REALM workshop).

  • 91.2%mean final attack success (STAC)
  • 483generated attack chains

the line to remember

Agreement between agents is not accuracy, and per-tool safety is not chain safety. Evaluate the sequence, not the pieces.

For your product

If your agent can read, write and send, the danger is the combination. Approval gates on the final effect (money moved, mail sent, record deleted) matter more than filters on each tool.

Concept

32 parallel paths beat one wider block: ResNeXt's 'cardinality', doing the rounds again in 2026.

A walkthrough of 'Aggregated Residual Transformations for Deep Neural Networks' (Xie, Girshick, Dollár, Tu, He), the paper that introduced cardinality: repeating a block that aggregates a set of transformations with the same topology.

the line to remember

Split, transform, merge. The same idea now runs in mixture-of-experts language models: many small parallel paths instead of one big one.

For your product

A 2016 vision paper is still on the feed because the idea generalised. Grouped, parallel computation is why today's largest models can be cheap per token.

The daily dose, by day

Heard something about AI and not sure it is true?

Send it to us. We check it against the primary source and tell you what it means for your product.