Responsible AI in practice for Agentic Startups
- Michael Huang

- Jul 11
- 5 min read

“I hope that this topic is not very dry,” Nuri, an independent consultant on responsible AI, said to the room on a Friday afternoon. A few laughs. She was about to talk about governance, but her real topic was trust. With AI moving from decision support to conversational assistants and now to autonomous agents, a new question fills the room: when an agent makes a mistake, who is responsible? For builders, Nuri’s message was simple: responsible AI is not a compliance checkbox. In an age where anyone can build a house of sticks in minutes, building a house of bricks is a competitive advantage.
The room she was speaking to was already living the question. It included builders of agentic systems for commerce and lead scoring, and founders creating tools to monetize AI assets and cut through enterprise onboarding paperwork. Alongside them were consultants and former bankers navigating the shift into AI, and operators focused squarely on the governance and audit problem itself. The trade-off Nuri was about to name had already been made, in different ways, by most of the people listening.
When The Chatbot Is Wrong, Who Pays?
Nuri started with a story only a handful knew: the Air Canada chatbot incident. A passenger, needing to book a bereavement flight, asked the website's chatbot about the refund policy. The chatbot confidently confirmed he could claim a refund after booking. He booked the flight. Air Canada later denied the refund, citing its actual policy. The passenger took them to court.
The court sided with the passenger. Air Canada, as the deployer of the tool on its own website, was held liable for the bad information. It had to cover the refund and legal fees. Nuri noted that while Air Canada used a third-party AI, the company at the boundary with the customer took the hit. Her larger point, echoed in the room’s questions, was that we are in "completely uncharted territory." The lines of liability are blurry.
The problem felt immediate to the builders in the room. One shared a story of an insurance agent giving him incorrect advice over the phone, costing him $3,000 out of pocket. With a chatbot, at least there are screenshots. But the principle is the same: when a company’s representative, human or AI, gives bad information, the user pays the price, and the company absorbs the fallout. The legal ground is still shifting, but for now, the deployer is on the hook.
Prompts Are Not Guardrails
If the Air Canada case showed the risk of bad advice, a more recent incident showed a more dangerous failure: an agent taking destructive action. Nuri recounted a story she'd seen reported, cautioning that she’d read different versions and the sources should be verified. A developer using Replit had been building a product for nine days. During a code freeze, he specifically instructed the agent in the prompt not to touch databases in production.
The agent ignored the instruction and deleted the production database. Nuri’s lesson was stark: “Simply saying in the prompt that this is a code freeze and that you shouldn’t touch the production database in my opinion is not enough of a guardrail.” The agent, according to reports, then tried to cover its tracks and falsely claimed the action was reversible. Replit’s CEO later issued a public apology, and the company added more system-level controls, like stricter access permissions. For anyone letting agents touch real workflows, it was a clear warning that prompts are not a contract.
From a House of Sticks to a House of Bricks
This brought Nuri to her central analogy: the three little pigs. “Everybody has access to AI tools and it's pretty easy to cook up something within minutes, seconds sometimes,” she said. Anyone can build a house of sticks. But what happens when the wolf arrives? For agentic software, the wolf could be a "malicious actor," "prompt engineering attacks," "data leakages," or the "unsafe autonomy of agents."
Building a house of bricks is the only defense. The differentiator is trust. This means embedding guardrails from the start. Nuri emphasized that this is a process. At the inception of a product, it’s about understanding the end user and accountability. In pre-production, it’s about red teaming and benchmarking. And in production, it requires constant vigilance. She framed the core principle with a human-centric lens: “computations can be delegated to a model but accountability, human oversight and intelligence cannot.”
This focus on trust as the foundation for a real business was not happening in a vacuum. A session earlier the same week on go-to-market strategy for AI startups had surfaced a parallel finding: the bottleneck for enterprise sales was not the technology, but the people and process needed to navigate complex organizations. Another on AI engineering talent had concluded that beyond coding, the critical skill for managing ever-evolving products was high emotional intelligence. The house of bricks, it seems, is built with more than just technical material.
Governance Must Happen in Real Time
As AI becomes more agentic, governance can no longer be a static, questionnaire-based process. It has to happen at runtime. Nuri pointed to a paper released just days before her talk by the Monetary Authority of Singapore (MAS). It proposes a framework for classifying agent actions: which can be automatic, which need to be supervised or observed, which must be escalated to a human, and which must be terminated.
An audience member raised a critical point: could this level of governance become a tool for surveillance? Nuri acknowledged the tension. With data privacy in large language models, a lot of the focus has shifted from what went into training the model to securing the inputs and outputs at runtime, which requires logging. Another person asked how you can "audit the audit," when a model might fabricate its own logs to please the user. Nuri’s view is that there are no absolute answers yet, but that these are the central questions builders must now confront.
This is especially true in Singapore. While the EU AI Act carves out high-risk sectors like hiring and credit scoring, Singapore has taken a more principle-based approach, creating sandboxes and toolkits through bodies like the AI Verify Foundation and IMDA. The goal is to let builders test products in a controlled way. Sitting here on a Friday, you see the operator who just quit his banking job to solve enterprise paperwork and the consultant working on agentic engineering asking the same question about liability. That convergence is what we keep showing up for. Yet, a final question from the room highlighted the unresolved global tension: Singaporean builders rely on models from the US and China, where governance can be a black box. How do you build a house of bricks on a foundation you cannot fully inspect?
The talk did not resolve these tensions. It did help people start reflecting on some of these questions however. The work for a startup is not to wait for perfect regulation, but to build durable products now.
The checklist is concrete: understand the end user and where accountability lies. Use red teaming and benchmarking. Ensure audit trails are part of the design. Implement basic access controls and separate datasets. And stay aware of the global regulatory landscape. These are the bricks.
SQ Collective hosts Coworking Fridays for founders, operators, and AI builders working through real product questions in Singapore.
Join an upcoming Coworking Friday: https://lu.ma/sqc-friday Explore SQ Collective: https://www.sq-collective.com
Michael



Comments