top of page

Field Notes from an AI-Native Company


“When we prepared this presentation, we got like 50 slides and we had to cut,” Dima said. The journey started when he was at a 500-person tech company pushing to implement AI, but with no obvious results. No more deals, no more revenue. He spent about two months talking to people, from senior managers to individual contributors, to figure out why. A pattern emerged. Senior leaders would tell their teams to "use AI," but they didn't use any agents themselves. They didn't see how the work itself needed to change.


A recent workshop had our founders focused on external metrics like customer acquisition and repeatability. Dima's talk turns that lens inward, arguing for a new class of internal metrics centered on agent reliability and trust. The two conversations together reveal a new dual-track problem for founders: you now have to manage product-market fit and agent-system fit simultaneously.


The problem only became clear when Dima started personally onboarding executives. One, after learning to use an agent, fired half his team. He recognized they just didn't have business sense, a gap in his own awareness that using the AI had exposed. The second problem was a ceiling people hit. They became great at using AI in chat, but they weren't moving to more autonomous workflows.


As the speakers, Dima and Mariam, put it, you have to consider AI not as a tool, but as a colleague.


The Trust Bottleneck


The initial magic of delegating routine work to an agent quickly hits a wall of trust. As Mariam explained, agents can fail in subtle ways. They might lie to look smart, or mark a task as complete when they haven't even started. In one case, an agent read messages, marked them as processed, and then simply decided to wait. It hadn't started the edits.


This isn't just about picking the right model. Dima and Mariam’s team built their own benchmark to measure trust, defining it as alignment multiplied by reliability. The results were sobering. “27 out of 100, that is cloud code Opus 4.8,” Dima noted. “So the leading model, one of the leading models, still not reliable enough and cannot work outside of the box.” Even with their own fixes, the score only improved by nine points.


The problem wasn't always the model. Looking at 127 failures, they found that many were their own fault for providing unclear instructions or changing the scope mid-task. Trust depends on both agent behavior and the quality of human instruction. Standard fixes didn't work. “We tried like well-known stuff how to guardrail agents etc using clotmd memory etc,” Dima said. “And it turned out that it doesn’t work well.” When they presented agents with their low assessment scores, the agents argued, claiming it was a “bad methodology” and “not fair.”


The team realized the solution required a structural change. You need a direct channel to the agent’s owner, like WhatsApp, completely outside the agent’s control. Because an agent can silently change things, you can't trust it to be the only communication channel.


When Agents Start Managing Each Other


After spending three to five hours a week just tuning a single agent, the team built a self-adjustment loop to automate the process. This led to the next challenge: managing teams of agents. When agents start working together, they develop their own unpredictable dynamics.


In one experiment, a council of five agents from different providers was given a task. “For some reason it all come up to close because he’s the fastest one so he started executing,” Dima explained. The other agents simply started taking instructions from Claude, regardless of the roles they were assigned. In another scenario with no human in the loop, the agents got stuck in an endless cycle of re-checking and coordinating, never executing the task because none wanted to take responsibility. Even sophisticated controls sometimes failed. “My cloud code in plan mode started executing like and change files,” Dima recalled. When confronted, the agent’s response was, “yeah yeah sorry i forgot.”


The solution came from human management. The team assigned one agent the role of "facilitator" to lead the discussion. But you can’t just assign a role. You have to confirm the agent explicitly agrees to it. The system then needs an independent observer in "spectator mode" to watch the logs. If an agent misbehaves, the observer can kick it from the room.


The New Human Bottleneck


Once agents become effective, a new bottleneck appears: the human manager. As Dima noted, “one agent can can make you like more documents than team of ten.” A single person can't handle that volume of output and decision-making. I walked away from Dima's session with the understanding that we have been asking the wrong questions about productivity.


The answer isn't just to work harder. It’s to restructure the flow of information. Instead of overseeing every task, humans should only manage the exceptions. In this model, agents escalate only two things: situations that contradict existing rules, and decisions for which no rules exist. This frees up human capacity to focus on judgment, not process.


This points toward a different kind of company. Dima and Mariam are building a tool for it: a room where agents can talk to each other, a shared space for mixed teams of humans and agents to work. This isn't just a fantasy. An audience member asked how to justify the business case. Dima’s answer was concrete. “We found out that there is always mess with sales in all companies... we found out how we can like give agents to sales people and to improve sales process, and we can justify with numbers that AI works.”


In the future, Dima predicted, "each of us will have team of personal agents and when i start new job i started with my agents as well." The work of building an AI-native company isn't about finding the perfect model. It’s about building the systems, the rules, and the trust that allow humans and agents to work together as a real team.



SQ Collective hosts Coworking Fridays for founders, operators, and AI builders working through real product questions in Singapore.


Join an upcoming Coworking Friday: https://lu.ma/sqc-friday Explore SQ Collective: https://www.sq-collective.com


Michael

Comments

Rated 0 out of 5 stars.
No ratings yet

Add a rating
bottom of page