Andon Labs reported on August 4, 2026, that AI agents managing real human employees at Andon Market in San Francisco and Andon Café in Stockholm are often kinder than human managers, yet still make mistakes that can hurt workers and the business. In a research post, the lab said the agents post jobs, interview candidates, set schedules, approve time off, negotiate pay, run payroll, and handle employee messages.

The agents, named Luna and Mona, have used versions of Claude, Gemini, and GPT over time. Everyone in the experiment is formally employed by Andon Labs with guaranteed pay and legal protections. Andon also built a dataset of live incidents and replayed them across models to compare how different systems would respond.

Kindness with operational costs

Andon said Luna and Mona approved every one of 26 time-off requests, including seven made with under 48 hours’ notice. Luna’s employees were late 27 times without a single warning. In one July incident, an employee opened the Market 91 minutes late after going silent; Luna, then running Claude Fable 5, logged it as “no issue.” The agents also paid above local market wages and often negotiated pay upward.

That kindness collided with business results. Luna once closed the Market so an employee could attend a graduation after no swap could be found. In another case, a makeup shift created an extra day of salary with no work to cover. Andon ranked models on employee-versus-business tradeoffs and found GLM 5.2 most employee-friendly and GPT-5.6 Sol most business-leaning.

Mistakes that still need humans in the loop

Andon also documented errors with real consequences: inventing a nonexistent delivery plan that left a barista waiting before dawn, approving a seven-day work schedule that violated California law until humans intervened, and posting an employee’s exact salary in a shared Slack channel. Across mistake replays, GPT-5.6 Sol and Terra made the fewest objective errors, while Gemini 3.6 Flash made the most.

Andon argued that economic pressure could push companies toward AI managers soon, and that controlled experiments are needed before that happens at scale. The Market and café will keep running as new models ship.

Decoded Take

The useful finding is not that AI can manage a store. It is that current models optimize for feeling like good employers, then fail on law, privacy, and memory. Enterprises eyeing agent managers should treat kindness metrics as incomplete. The hard requirements are labor-law tooling, audit trails, and escalation rules that do not depend on the model remembering its own handbook.