0:00
/

How Cisco's Outshift Runs Engineering On AI Tiny Teams And A Memory Engine

Vijoy Pandey shrank his teams to five people and explains what that costs, and the memory engine and decision leads Outshift built to hold it together

Most conversations about AI and team size stop at the headcount. Teams get smaller, agents absorb the work, and the argument ends there. Vijoy Pandey, SVP and GM of Outshift by Cisco, joined the Deep Engineering Podcast to argue that the problem starts immediately after that, because a company running a hundred small teams still has to make them add up to one company.

He writes about this work at Coherent Cognition, his Substack on scaling out infrastructure for non-deterministic systems, and posts regularly on LinkedIn.

For eight months Pandey has run Outshift on a model he calls T3, tiny teams with tokens, where the delivery unit is one to five engineers with heavy agent support. His argument is that a business needs three systems to produce outcomes, a productivity unit, a context substrate and a coordination mechanism, and that AI has collapsed the first while leaving the third almost untouched.

So Outshift built its own organizational memory engine, TOME, which holds designs, decisions and escalations for every team, weights chats and meetings above documents, and keeps anything an agent writes as a permanently secondary source. Pandey is also direct about the limits, since the coordination tax arrived at around five teams and none of this structure earns its cost below roughly thirty people.

Below is the cleaned up transcript of our conversation.


Tell us what Outshift is inside Cisco, and what you have changed about how your teams are set up and how they operate day to day

Outshift is Cisco’s incubation engine. We look at emerging technologies, and right now that primarily means agentic computing and quantum computing. We put the Cisco hat on and figure out what products Cisco can build in those spaces. We build those products, get some customer traction, get some practitioner adoption, and in parallel we work with the business units to scale the product and the business out as part of larger Cisco. Our job is to reduce risk for Cisco when it enters an emerging technology space and that market.

On what is happening operationally, your readership will already be aware of the trend. AI has made building things dramatically cheaper, and that means coordination is becoming the part that slows companies down. A business needs three systems to function and to drive business outcomes. It needs a productivity unit, a context substrate, and a coordination mechanism. That was true regardless of AI, 70 years ago and 100 years ago.

What is happening now is that the first one, the productivity unit, is shrinking massively because of agentic AI. On context, a whole set of startups are coming up to build context platforms and memory platforms. So everyone is racing to build the agentic productivity unit, everybody is building and buying the context layer, and nobody is talking about the third pillar, which is how hundreds of these tiny teams talk to each other and add up at the company level to drive a business outcome. Context is not coordination. What we have built, which we call the T3 operating model, is our attempt to solve all three pillars so we can accelerate business outcomes in a world of agentic native development.

A team of one to five developers with a lot of agent capacity behind it is hard to picture from outside. What does a T3 actually consist of

Let me start with why teams are even shrinking. If you have heard the recent chatter, especially the commentary from the foundation model labs and some of the AI native forward companies, including a recent one from Coinbase CEO Brian Armstrong, what they have been saying is that you are getting towards a one person company. Brian made the comment that you are moving towards a one person team. I would take that as directionally correct but not absolutely correct. Brian and others are right that teams are shrinking. The real test is whether hundreds of these can operate as one single company.

There is a history here. Teams have been shrinking for the past few decades. If you go back to hard engineering, the way you make cars, the way you make rockets, and in our own business the way you make hardware switches and routers, you are not shrinking teams as much. You need large specialist teams for cars and rockets and routers and switches, because every person brings something different to that equation.

With the advent of cloud and cloud services and the entire API model, a lot of that shifted towards what we now know as two pizza teams. That is what Amazon and Jeff Bezos and the crew were proponents of, and it has been deployed everywhere now. Cloud moved us from really large hard engineering teams to two pizza sized teams, eight to ten to twelve people in size, and that became the size of the delivery unit.

But the coordination problem became harder, because coordination moved from intra team coordination in large hard engineering teams to inter team coordination in these two pizza teams. A hundred people in one team is intra team coordination. A hundred people in ten teams of ten is inter team coordination. The way Bezos and team solved it is through what is now known as the API manifesto, where every team coordinates with the team building a service through APIs, not through design talks and not through meetings. That codified coordination and made it simpler, and it solved inter team coordination.

That is happening again with AI, where team sizes are moving towards one to five. The coordination surface is exploding again, and APIs are not going to be sufficient to solve it. That is what we are trying to do with our model. The first thing to realize is that it is not the size of the team, it is the fact that the teams are shrinking. It is not the agentic tooling you are using, and you can use many types of agentic harnesses and tooling. It is that artifact generation is simpler, and because of that teams are expanding in scope while shrinking in size. That leads to coordination problems, especially inter team coordination, and the big thing to solve is that coordination problem.

You started with one team on one outcome, then took it across the whole organization. What was in the model by the time you rolled it out that was not in that first team

We are eight months in, we have looked at every type of metric, and we have looked at what worked and what did not. There are a few things we have learned already, and we are still learning, because we are very early in this journey.

First, the architecture of what the three pillars look like. T3s, tiny teams with tokens, are our productivity unit. They are our artifact generation unit, and they are human forward, human led teams assisted by many agents. So multi-agent teams with human centric decision making, judgment and accountability.

For context we built TOME, the organizational memory engine. It is the memory engine for the entire organization and it is a layer we built from scratch. Every team stores and retrieves context there, everything from design to decisions to outcomes and how you measure outcomes, and conflicts and escalations. That last part matters.

The coordination layer is more of a process and a structure at this point. Internally we use objectives and key results. My background is from Google, which has used OKRs for a long time, and it is normal in cloud centric teams, so we inherited OKRs as the way we measure ourselves regardless of the agentic world we live in.

For coordination we broke teams, artifact generation and leadership into three buckets. Objectives, of which a team or an organization might carry three to five. T3s, the productivity unit. Objectives do not map cleanly onto T3s, that is too much of a jump, so we have something in between called areas. You can think of areas as functional areas, as architectural alignments, or as key results, and you decide what works for you. We appoint context leaders at both levels, so objective leads and area leads, and their entire job is to keep things consistent, to make sure decision making happens at the lowest layer possible, and to make sure escalations move fast. All of that happens through TOME.

On learnings, the first is that not all sources feeding into TOME are equal. People usually say recency bias is a bad thing. For TOME it is a good thing. TOME prioritizes chats and meetings over everything else, because that is the equivalent of a hallway conversation, and those are the most relevant and the most recent.

We learned that the hard way, because of the second thing we learned. Every team, including ours, has a very large collaboration surface. We have Git and GitHub for code. We have the whole Atlassian suite, Jira and Confluence, for ticket tracking and documentation. We have Office 365, because we build presentations and Word documents to share with other people. We have video meetings and transcripts, Slack chats, emails. The collaboration surface is massive, and a surface like that is primed for authoritative source failures. Which source do you trust. At that point you have a data architecture problem, and it is unmaintainable at scale. That is why we built an architecture that prioritizes, and hallway conversations, meaning chats and meetings, get priority.

The third learning is about agents. Cisco has rolled agents out across its entire workforce of 90,000 people. We all have personal agents, and I am running [tool names, confirm with Vijoy] and a number of other things. Having agents, or model harnesses like [Claude Code and Cowork, confirm], that can plug into the whole collaboration surface does not mean you no longer need a context layer like TOME. A context layer does not just plug into those surfaces, it has authority and hierarchy built in as first class principles. TOME drives authority, hierarchy and consistency across all of them.

One design goal we hold is that at some point I want to get rid of all these collaboration surfaces. Today we go from people to a massive surface to TOME. I want to flip that, so we go from people to TOME, and then to a collaboration surface only when that surface is needed, for outbound artifacts like a presentation or a document or a PDF going to somebody else. That is a challenge to the team and it is what we want to get to.

The case engineers keep running into is not looking up a fact, it is looking up an argument. A team formed last month needs to know why a decision was made six months ago by a team that has since dissolved. How does that resolve at Outshift today

Human judgment is critical, and it is even more critical in this agentic era, because artifact generation has become cheap, simple and plentiful. What is happening is not just that my work is getting faster with agent assistance. My work is expanding. My role expands beyond what I could do yesterday. If I am an engineer, I could write code yesterday. Today I can write code faster and at a higher abstraction layer, but I can also work out what the market needs are. I can work out what a good design or user experience might look like. I can work out what customer needs are and what will resonate with customers. These tools let me expand my role into aspects further down from it, and that makes the coordination layer harder and human judgment more important.

Through TOME we apply the Bezos philosophy of how you prioritize decision making, which says some decisions are one way doors and some are reversible. You make a one way decision and you cannot come back to it, because coming back is expensive. Others are revolving doors, where you decide, move on, and the cost of reverting or pivoting is low, so you should make that decision quickly.

When T3s are formed we work out whether the outcomes they generate are one way or reversible. One way outcomes get more scrutiny and more time. For reversible outcomes we use T3s and TOME and go forward, so ship, iterate, do not worry about it.

Second, we look at decisions that carry dependencies. Some T3s generate outcomes that other T3s depend on, and the cost of those decisions is higher than for outcomes that are independent of each other.

The third is blast radius. If an outcome from a T3 goes haywire, how bad is that for the team, the product and the company. Based on the blast radius, that decision holds weight or it does not. Those are the things we use to work out how critical a judgment or a decision is.

What TOME then lets us do is hold that decision matrix, let the objective leads and area leads make the right calls at the right layers, and run escalations asynchronously and quickly according to the weight of the decision. One test I set the team is what I call the 12 a.m. test. I wake up at 12 a.m. on a Sunday morning and want to understand the impact of a business outcome. Can I get the what, the why and the how of the decision behind it, at 12 a.m. on a Sunday, without bothering anybody else. If I can do that, TOME is successful. We have managed to get to that point.

Once agents can generate documents at almost no cost, a store that accepts everything fills with plausible material nobody checked. What are you deliberately keeping out

This goes back to how we constructed TOME. The collaboration surfaces are the prime source of hallucinations, and that was a lesson we learned when we started out.

We designed TOME as the layer after the collaboration surfaces. You have people augmented with agents in the shape of T3s, one to five in size, still talking to our collaboration surfaces, because that is what we are all used to. We write in Confluence, we write in Word, we generate PowerPoint, ops people take meeting minutes, emails go out, there are calls and chats. If you feed all of that into TOME, which is where we started, the authority, the accountability and the data architecture fail you, because the tool will hallucinate. You can throw data at TOME and at agents, and TOME is agentic in nature, but it will try to give you its best answer. You do not want its best answer. You want the right answer.

What we have done since is build an architecture that says these are the primary sources, the sources of truth, and these are the secondary sources. In the primary sources, humans are the authority. In the secondary sources, humans and agents are the authority. So when agents write design documents and presentations, and when they collaborate with us in chats and meetings, they are always secondary sources. The primary sources of truth today are still human maintained and human constructed, by that context leader layer, which is why the third pillar matters so much. The objective leads and the area leads are the ones who maintain that human source of truth. That is what we do today. We hope to change it, but that will take longer.

How is the boundary between two teams actually expressed. Is it a written contract, a shared artifact, a person, something else

Right now the boundary between teams is expressed through the context and coordination layers. Teams are T3s, one to five people, heavily agent augmented. The context layer is where we store decisions, outcomes, what we measure, escalations and so on. Coordination, instead of being codified, is human led today. Context leaders at the objective and area levels make sure the sources of truth are maintained and that what comes out of TOME is not a hallucination exercise. That is where we are, because we are rolling this out slowly and there are many collaboration surfaces.

You are right that as engineers we think about this as a software problem. Right now T3s are human led and agent assisted. We are moving towards a world where agents and humans hold the same level of agency, so multi-agent human teams where a human has the same authority and agency as an agent. In that world, which is not far away, how do you codify the coordination pillar.

What we are building for that is what we call the Internet of Cognition, which is Outshift’s own project at Cisco. It is the coordination layer, and it lets agents and humans coordinate towards a common goal, define common intent, and define a negotiation mechanism through protocols. We do that all day long through language, and the language of machines and agents is protocols. So can protocols carry meaning rather than only data. Once they carry meaning they can coordinate, they can negotiate, and they can reach grounding on terminology.

These are things we do all day. Half the problems we see day to day, even with T3s, come down to whether I mean the same thing as you, whether we are aligned on the same terms, whether we are aligned on the same KPIs. That is where half the escalations happen and where half the friction in the seams happens. Grounding, coordination, negotiation and evaluation all derive from meaning. If we can build protocols that carry meaning, we can codify the coordination layer and move towards multi-agent human teams doing shared reasoning together to solve for business outcomes.

A shared memory system can show everyone that a conflict exists but cannot pick a winner. Where is that handoff in your setup, and what does it look like for the service owner on the receiving end

TOME is the context layer that surfaces these disagreements. It surfaces all kinds of outcomes, successes, failures and conflicts. The context leaders at the objective and area layers are the ones who resolve them, and they do it asynchronously, as soon as they see it surfaced.

TOME surfaces the disagreement. If two T3s cannot resolve it, TOME reflects that. If the area lead cannot resolve it, TOME reflects that. If the objective leads cannot resolve it, TOME reflects that. At that point it comes to me and my senior leadership team.

We still run standup meetings, and what we do in them is walk through the escalations the objective leads could not resolve through TOME, go through them one by one, and take the decision. So the standups now happen at the objective layer, which is where they should happen, because objectives are what you should measure. You should not measure productivity. You should not measure the number of lines of code you wrote. You should measure whether it solved a customer problem, whether practitioners are using what you built, and how quickly that outcome is happening. Outcome velocity.

So yes, there is still decision making with humans in the loop, but it happens at the objective layer, at the business outcome layer, and it happens through TOME asynchronously and with extreme velocity. We are not taking humans out of the loop. We are making sure artifact generation happens with velocity, through T3s and agentic software development, and that decision making and coordination also happen with velocity, through TOME and this layer we have built.

All three pillars have to exist eventually. Which one would you build first, and what does it cost to get that order wrong

At team sizes of one through thirty you do not need this structure. If you are a small startup under thirty people, you can use agentic harnesses, build your T3s, and end up with five to ten of them. You or a small group can be the context and coordination layer yourselves, very human centric. So your productivity pillar can be agent driven and agent forward from the start, and the other two pillars can be one person or a few people. You will not hit velocity problems, because there is not enough context to keep and you have maybe five to ten teams.

Beyond thirty people, give or take, is where we see the explosion take shape, because of where we started. As you shrink team sizes, coordination shifts from intra team to inter team. As you go from five or ten teams towards thirty teams and towards a hundred teams, your coordination problem is an n squared problem, and n squared is exponential in the number of teams rather than the size of the team. So from ten or more teams onward you start hitting velocity problems, not in artifact generation but in context maintainership and decision making.

That is the pivot where you should deploy a context layer. You should break your objectives into areas and T3s, and appoint context leads whose entire job is to keep context sane and coherent and who hold the authority to make the call at the lowest layer. Organizational design is actually decision making design, and that is important.

How do I know that. When we started moving to the T3 operating model we started with one team, and in all things like this you solve it recursively. We deployed T3 on the T3 team itself, so the team building TOME used TOME to deliver TOME. It is a very meta and very recursive problem. It was one team and we tried it out. Then we rolled it out to three teams, five teams, and eventually to the entirety of Outshift. We saw the model start breaking at around the five T3 mark, because that is where the coordination and decision making tax started hitting people, and we needed infrastructure in place.

What would you advise engineering leaders that most people do not know

Agentic development is changing the game for everyone. It expands what any individual can do into many more things. But there are human values that are hard to replicate in agentic development. One is judgment. Another is influence. The third is accountability.

Can you make the right calls with a lack of data. Agents do not do that well. Can you influence other teams and other people to work with you on the vision you have and the decision you made. Influence is a big one. And once you have done those two, are you the one accountable, because you will be accountable for the outcomes that result. Taste is the other one that is very human.

So when you build these teams, use agentic development, but right size the teams because of judgment, taste, influence and accountability. That is how you end up at one to five people. Then build out the context layer and the coordination structure, because without those you will be in decision making hell. You will feel you are shipping fast, and you will be shipping artifacts really fast, but you will not be shipping outcomes fast, and that is what every business is after.

Discussion about this video

User's avatar

Ready for more?