A single LLM call only gets you so far.
On its own, a prompt in a chat window is isolated, prone to hallucinations, and trapped on one person's screen. It may save time. It may also generate plausible work that someone still has to verify and clean up.
The larger value does not come from writing a better prompt. It comes from stacking LLM calls into an architecture: specialized agents, clear hierarchies, autonomous execution loops, shared tools, and human approval at the right points.
When those parts have clear boundaries and checks, individual speed gains become a capability the whole organization can use. This is how our team at BRACKETS moved from chat windows to a hybrid AI-human team.
Stage 1: AI moved into the terminal
We started where most teams did. We used LLMs as code completion tools, chat assistants, and a quick way to analyze text or summarize documents.
That changed around late 2025 and early 2026, when stronger reasoning models arrived. We moved much of our work from ChatGPT to Claude and started using Claude Code extensively in the terminal. It became our standard engineering interface, while the web interface remained useful for other work.
The terminal changed the role of the model. It could read a whole repository, execute commands, run tests, and modify files directly. It was no longer answering questions beside the work. It was operating inside the work.
Non-technical colleagues soon saw the same potential. The interface could support recurring research, market reporting, and internal workflows. Tools such as Claude Cowork made those capabilities available without asking everyone to work in a terminal.
Stage 2: Tools and specialized agents
Model Context Protocol gave models a standard way to use tools. Calendars, email, CRMs, and databases became part of the workflow.
Instead of manually moving information between systems, a colleague could ask the model to analyze a meeting transcript, save the notes to the CRM, and schedule a follow-up. The model could complete the sequence through connected tools.
At the same time, we began creating specialized agents. Each one had a defined role, a concrete set of skills, the context for its project, and limited tool permissions. An engineering agent could be tailored to one framework and repository. A sales agent could work with research and CRM data.
The agents could also delegate. A coordinator could receive a broad goal and assign focused tasks to agents for research, copy, design briefs, or implementation. Context no longer had to fit inside one large prompt. It could be split according to responsibility.
This is one of the practical differences between a chat assistant and agentic engineering. The work is decomposed, permissions are explicit, and every stage has an owner.
Stage 3: Autonomous loops in cloud sandboxes
The next constraint was the synchronous loop: prompt, wait, review, prompt again. We replaced much of that hand-holding with structured autonomous execution.
An agent could receive a backlog of 30 or 40 tasks, process them in sequence or in parallel where the work was independent, run tests, correct syntax errors, and continue while the human team worked on something else.
Long-running work did not belong on a laptop. Sleep mode, limited compute, and access to local files all made that setup fragile. We moved execution into isolated cloud sandboxes built for each project. The agents could continue running, while their access stayed inside the environment they needed.
The architecture supported complete workflows. One process turned a rough brief into granular user stories for human review. Another took the approved stories through implementation, testing, and security checks.
We used that workflow for a one-to-one legacy codebase migration. The project took two weeks in total:
| Phase | Duration | Responsibility |
|---|---|---|
| Preparation and rules | 4–5 days | Human engineers defined the flow, target architecture, prompts, and test guardrails. |
| Autonomous execution | 2.5 days | Agents rebuilt the application and ran tests through thousands of sequential LLM calls. |
| Review and polish | 5–6 days | Human engineers tested edge cases, worked through corrections, and completed the security review. |
A single prompt could not have rewritten the application reliably. The result came from the structure around the model and from the human work before and after autonomous execution.
Stage 4: The work became visible in Slack
The technical workflow improved, but collaboration remained difficult. AI sessions still lived on individual machines. Colleagues could not see what an agent was doing, add context, or make a small decision without asking the person running the session.
Claude Tag moved those interactions into Slack. A task could start in a public channel, an agent could pick it up, and the team could follow the work and correct its course in the same thread. The shared context removed many status updates and handovers.
The economics were less useful. Channel activity was billed per token rather than through individual subscriptions. At team scale, the projected cost reached thousands of dollars per month. The pricing discouraged the shared behavior that made the setup valuable.
Stage 5: A provider-independent agent layer
We wanted the shared workflow without tying the whole system to one model provider. We moved the agents to AgentConnect, an open, provider-independent platform. The tools, skills, and boundaries could remain in place while the underlying model changed through configuration.
That gave us specialized Slack agents for several kinds of work:
- Engineering: agents tailored to specific stacks, frameworks, and repositories.
- Marketing and sales: agents handling research, campaign preparation, and CRM updates.
- Legal: agents checking agreements against defined rules.
- Operations: general agents supporting recurring internal work.
Each agent has explicit access to projects, tools, and Slack channels. The rules also define which agents may communicate with one another.
That allows a support workflow to split naturally. One agent analyzes a bug report. Another drafts a response for a colleague to review. A third starts the fix in the background. For a feature request, agents can ask the requester for clarification in Slack, prepare a pull request, run peer review, and hand the result to a human for final approval.
The process keeps the software engineering controls we already rely on: code review, automated tests, security checks, and human sign-off. The architecture changes who performs each step and how quickly the work moves between them.
Where we are today
We now have 17 specialized agents operating across our workspace and processing roughly 60 million tokens each day. We are still testing the limits, adding capabilities, and tightening the boundaries when we find a weakness.
In a single Slack thread, three people and two agents may work through the same feature. Everyone can see who owns the decision, which system each agent may access, and where human approval is required.
That visibility matters as much as model performance. A private assistant can make one person faster. A governed agent architecture can make the work reusable, reviewable, and available to the whole team.
We did not replace the team with AI. We designed a system in which people and agents carry different responsibilities. The value is in that architecture.