Reading 66

Some thoughts on Agent Ecosystems

R.V. Guha, 2026

Summary  

I make the case for an ecosystem of millions of agents*, any number of whom can be found and invoked for a given user task --- more like the web than like smartphone apps wherein an agent would be limited to a small number of ‘installed’ tools. This activation of the torso/tail will benefit Foundry and Copilot Studio, as the platforms for building them, M365 Copilot as the entry point into them, and a new ‘Agent Finder’ as a service for finding them.  

 

Information Pipelines    

Agents need information from the outside world. And with the rise of reasoning models, the amount of knowledge these models can effectively bring to bear on a problem is rapidly rising. The effectiveness of agents will depend as much on the information it can access as on the underlying model.  

 

Chatbot applications typically run the user’s prompt/query (or some derived fanout) over personal/enterprise document / database corpora (documents stored in SharePoint, Google docs, etc.) and augment the LLM context with this retrieved content.  

 

However, not all the information needed is in these content stores — a lot of it resides in the external Web. Chatbots started getting this information by making calls to web search engines. However, (as described in a later section) keyword-based web search is both a very thin and fragile straw through which to feed knowledge hungry LLMs.  

 

To solve this problem, OpenAI, Gemini and others have started augmenting results from search with ‘side corpora or databases’ with different kinds of data. We are seeing an increasing number of data licensing deals being done for this data. It is also notable that OAI’s deal terms are sometimes exclusive, i.e., the data may only be licensed to them.  

 

This ‘deal based’ approach, which is expedient, has shortcomings. A company can realistically only do deals with a limited number of partners. As the Web and Web search showed, there is a very long torso and tail and there is a significant demand for knowledge in that torso/tail. We need a platform/ecosystem based approach.  

 

We have seen many iterations of the competition between deal-based inclusion of high value content versus ecosystems that lean into the torso/tail. Two notable examples are  i) AOL (and other proprietary services) vs the Web ii) Yahoo (and other portals) vs Google search.  

 

However, under the right circumstances, the ‘deal’ based approach sometimes wins, as we have seen with music, legal databases, etc., where the gravity of the data creates a moat. Examples in ecommerce include Amazon marketplace, Ebay, Airbnb, and others. Of course, each of these is slightly different, but they all have the property that the popularity of the market leader makes it easier for them to obtain deals which in turn strengthens their position. In the chatbot space, content providers want to show up on ChatGPT and are doing deals with them. Though OpenAI’s focus has been in the consumer space, they are now leveraging consumer strength to push into enterprise. They could have adopted a more torso/tail friendly approach but have chosen not to.  

 

The deal-focused approach works for up to hundreds or maybe thousands of sources. It does not work for millions of sources.  Creating an ecosystem of millions of sources, that can be effectively tapped by Copilot, will allow us to differentiate.  

 

As mentioned earlier, chatbots already use search. We can use many of the ideas in search, while addressing their limitations for this purpose. The core limitations of search, as conduits for external knowledge into LLMs are:  

  1. Agents will want to ask complex questions that can be handled by today’s LLMs, but are beyond the scope of today’s keyword search engines. In contrast, agentic interfaces have already shown that they can answer much more complex queries.  
  1. Web search inherits the limitations of its crawl — freshness and completeness. Most indices today have 0.5-2 % of the pages crawled. Pages in the index are often months stale. The crawl-index model of search was developed 25 years ago, before hyper scalers, when web latencies were much higher and the web was much smaller.  

Imagine if we were to design a system for feeding information from the Web to LLMs. Would it really look like the 10 blue links that were designed for human consumption about 30 years ago? In addition to the limitations mentioned above, we need something that supports actions.  

 

The proposal is to help create an ecosystem of agents, who can interact in common protocols and formats (NLWeb). In the limit, we want every website and every application to have an agentic front end that can respond to rich natural language queries. Given a query or user intent, a Chatbot can, during inference, ask the right set of agents, in natural language.  

 

While this might appear laughably daunting, it is not so. We already have Cloudflare, Tollbit and others enabling their customers to create agentic endpoints that speak NLWeb. At Ignite we are on track to make it easy for anyone running Elastic, Yoast or Wix to create, with one click, an NLWeb agent. With luck we will have a few others, like Shopify. Some of them will use our code, some won’t, but since they will all speak the same protocol, an agent talking to them should not know the difference.  

 

The real value will come from small companies like tinyfish.ai . TinyFish builds ‘Nano-Agents’ that use a combination of techniques to create a natural language interface to small sites. They have created 20k+ nano agents and were able to make these NLWeb in a few hours.  We are partnering with them and Tripadvisor so that the 10k+ tour guides who have a page on Tripadvisor will soon each have their agent available to speak to Tripadvisor’s users on that page. And Tripadvisor can aggregate these nano-agents. So even a small, locally-owned walking tour site in Italy (Walksofitaly.com) can instantly answer not just “what do you offer?” but also “when is the next availability, and how much will it cost?” Instead of manually navigating dozens of listings, a user could simply ask, “which walking tours in Italy are run by locals, available this weekend, and under €50?”, and get the answer.  

 

Agent Finder  

In a world of millions of agents, we need a mechanism for our Chatbots to find the right agents to help with a given task.  This will have similarities to a search engine, but since this will be an index of agents/sites, not pages, both the size and churn will be many orders of magnitude smaller than for a search engine. We want an ecosystem with multiple Agent Finders. Like Web search, the index should be publicly available.  

We also borrow from DNS, allowing for cascades of Agent Finders (e.g., one for agents behind the firewall, which aggregates with one for agents on the public Internet). Agent Finder uses the same protocols and formats that NLWeb does. We are designing the discovery mechanism for the Finders to find agents.  

 

Unlike Web search where the human user has to click through to each page in the results, the Agent Finder has an option where it can query the agents, aggregate the results and return them to the calling agent.  

 

We have built an Agent Finder that we hope to launch as a public service in the Ignite timeframe.    

 

Structure of Agent Ecosystems / Distributing complexity  

There is a lot of excitement around MCP (and competitors like A2A /UTCP/ANP) to expose data and services to Chatbots. While this layer is important, it allows for many different control structures in the Agent ecosystem. The MCP layer is like HTTP. HTTP could have been used to expose the APIs of different content management systems (CMS) with the expectation that the browser has to understand the API of a publisher’s CMS for a user to be able to view that publisher’s content. Or, we could (as we did), introduce another abstraction layer (HTML+Javascript+CSS). MCP/A2A/… allow each website/application to expose an arbitrary set of tools and unsurprisingly, each of them is exposing the internal APIs to their content management system. We could continue down this path or introduce the abstraction layer — which is what NLWeb tries to do. The decisions made in the next few months will determine both how much the ecosystem will scale and how we are likely to ‘apex’ aggregators.  

 

At its core, the issue is whether we have a small number of super smart agents that invoke everything else as a dumb tool or we have a large number of smart agents that invoke each other to perform tasks. And this will depend on what goes on top of MCP/A2A/UTCP/ANP, not on the choice of which of these takes off. In other words, is there a very small number of standard tool end points (very loose analog of GET and POST — ASK and DO with NLWeb) with standard return structures or a wild west where every MCP endpoint has a different set of tools that the calling agent has to understand.  

 

We illustrate the core issue with an example. Imagine a chatbot at a company, say, in the context of marketing some safety product, needs to determine the number of unemployed, number of households making more than $100k/year and the number of thefts in a given city. This data is available at BLS, Census and FBI. The arguments to the APIs require knowledge of the fairly arcane and different ways of naming and referring to places (FIPs codes, census variables, FBI place names, etc.) To further complicate things, new places keep appearing and variable semantics change.  

 

Government data sources are not unique in the complexity involved in identifiers, variables, schemas, etc. Any medium sized etailer with even O(50k) SKUs has this level of complexity. Real Estate, Finance, etc. each have their own dimensions of complexity.  

 

We want a system that goes from a user intent expressed in natural language to this data. Something(s) somewhere has to deal with this complexity. The question is — is it all in one place or is it distributed.  

 

There is one school of thought that only asks the content/service providers to expose their existing APIs via MCP/A2A/… and have the model take care of the mapping from natural language to the APIs. This centralizes the complexity. Let's call this the centralized intelligence (CI) model.  

The other school of thought asks the content/service providers to each abstract their complexity so that they each take care of the NL to API mapping.  The Chatbot then simply uses the NL query to the service provider, via MCP/A2A/Http/… This distributes the complexity and leads to a ‘community of intelligence’ (COI) model.  

 

The COI model is closely tied to the proposal in the first section to create a web of potentially millions of agents, speaking natural language. The COI model requires an additional service — something loosely similar to a search engine, which given a user intent, can tell a Chatbot which agents it can ask for help. This is the ‘Agent Finder’ referred to in the last section.  

 

At one level, CI vs COI is simply a matter of ‘which’ tool we expose via MCP/A2A/… but the two approaches lead to very different kinds of ecosystems.  

  1. Master-slave vs peer: CI tends towards the master slave where the central model largely determines whose data is used and how it is used. COI on the other hand should allow for more autonomy to every service provider.  
  1. Apex aggregator: The last two decades have been characterized by the rise and subsequent domination of the Internet by ‘apex’ aggregators (Yahoo, Google, Meta, et. al.). CI will lead to an even greater concentration of control.  
  1. Scaling: Progress in tool use has been slow. Even though the concept has been around for nearly 4 years (since WebGPT), even the biggest models today struggle with tool selection and invocation when given even a few hundred tools. They also struggle with tools whose ‘function signature’ is complex (like in the government data above). This has implications, not just for scaling, but also for business arrangements. If a CI can only handle a few thousand (and that is being generous) tools/services, many of which require custom post training, this will lead to CI builders carefully curating the set of tools. This in turn leads to a small number of bigger players, ignoring the torso and tail. This is very unlikely to reach Web scale.  
  1. App install vs web search: The smartphone app ecosystem is very different from the web ecosystem. The former has strict gatekeepers, and a typical user has fewer than a hundred apps and the apps they use are restricted to those. This is of course in sharp contrast to the web. CI tends to a smartphone like ecosystem. COI is more web like.  

 

One entry point versus many  

There is a natural tendency for a given user to start everything in one place and that will not change. That starting point has been a combination of browser/search engine – Netscape to Yahoo to Google to .... There is a natural tendency for this apex aggregator to devolve into a walled garden once it has achieved dominance.  

 

One of our design goals this time is to create an architecture that is more resilient to this. Today, creating a search engine has huge upfront costs --- crawl the web, keep it updated, get user data to fine tune, etc. Federation was not realistic at the time today’s search engines were developed. But when vastly reduced latencies and hyper-scaler induced concentration of servers, it is now a possibility. It is this change in the technical infrastructure that Agent Finder leverages. And since the index for an Agent Finder is order of magnitude number of sites, not number of pages, something that is both much smaller and changes much less rapidly, the cost of creating a new aggregator should come down dramatically, thereby making it easier for new entrants.  

 

 There are other factors that might make it difficult for users to switch between aggregators. Foremost amongst these is conversation history and memory. If we make it easy for anyone creating an agent or aggregator to store conversation history, we can both stay relevant irrespective of the aggregator and also ourselves provide services on top of this conversation history / memory.  We should also provide a Foundry service that makes it easy for any agent creator to use Foundry to store conversation history.  

 

↑ Top