Case study 02 · KaDeep Technologies
TestStudios — Testing becomes a conversation
A chat-first QA agent that tests web and mobile apps for you.
- Product
- TestStudios
- Maker
- KaDeep Technologies
- Category
- Developer tools / QA
- Platform
- macOS desktop (Electron)
- Version
- Shipped on skills-first architecture
- Website
- kadeep.ai
90-second brief
How to read this
The low-code IDE still assumes someone will record, edit, and maintain a suite. Non-technical testers, PMs, and support still bounce off a script-shaped product. Web and mobile still live in different tools. I needed a Studios that removes the scripting layer, not another nicer editor for it.
Keep low-code as the only Studios
Died
Record-and-heal is a real product. It is still a barrier. We would keep selling to SDETs and losing the people who file the bugs.
Fully autonomous agent, no plan
Died
An agent that clicks before a human sees the plan is a liability in QA. Wrong path, wrong env, burned credentials. Fast demo. Unshippable for a release gate.
Chat-first, Plan mode, HITL
Shipped
Conversation is the authoring surface. The agent proposes a plan, waits, then drives a real browser or device. Pause and take over stay first-class. This is what shipped in 0.1.0.
- Decision
- Remove the scripting layer. Testing is a conversation with a plan you can approve. Web and mobile share one desktop app. Memory of the product is the category claim — not a one-off chat that forgets.
- Result I will defend
- Shipped on the skills-first stack: English in, structured bug report out. Chat, Plan mode, live view, OTP. Web and mobile in one macOS app. Weekly authors are not on a dashboard — I will not invent a seat count.
- What I got wrong
- I almost treated this as a skin on the low-code IDE. It is a different ICP. If the first screen is still a test table, the no-code bet is already lost. I also cannot pretend cloud device emulation is a device lab — Appetize is how we ship mobile without hardware. That constraint has to stay visible.
- Who loses if this is wrong
- The release manager, if an unreviewed agent path papers a defect. The tester, if chat cannot finish login. The SDET, if no-code floods suites nobody can debug.
Metric spine
- North star
- Time-to-first-executed-test
- A PM or QA can describe a flow and see the agent driving a real browser or device without writing a locator. If we lose to “just ask ChatGPT and click along,” we lose the wedge.
- Guardrail
- Unreviewed action / silent drive
- The agent must not burn a production login or a paid environment because Plan mode was skipped. Pause-before-act is a product requirement, not a preference.
- Not instrumented
- Weekly active authors who never open a script
- That is the ICP proof. 0.1.0 does not have a published cohort. I will not invent one.
- Not instrumented
- Plan-accept rate vs takeover rate
- Would show whether Plan mode is load-bearing or theatre. Not on a dashboard yet.
- Not instrumented
- Web vs mobile run mix
- The category claim is both platforms in one tool. Split is not instrumented.
Remove the scripting layer. Conversation is the test. Plan, then drive — web and mobile in one app.
Low-code still asks someone to think in steps and locators. A fully unsupervised agent is a demo. The shippable product is chat-first authoring with a plan you can reject, a live view you can trust, and memory so the agent is testing this product, not a generic website. That is a different Studios from the recorder IDE — on purpose.
Describe
A test, a bug, or a plan in English.
Plan
Steps for approval. HITL can reject or edit.
Drive + remember
Live browser or device. Pause, takeover, persist context.
