Ridhi Jolly
Back to work

Case study 02 · KaDeep Technologies

TestStudios — Testing becomes a conversation

A chat-first QA agent that tests web and mobile apps for you.

Product
TestStudios
Maker
KaDeep Technologies
Category
Developer tools / QA
Platform
macOS desktop (Electron)
Version
Shipped on skills-first architecture
Website
kadeep.ai
Create surface — chat-first. Projects in the sidebar, Plan mode under the composer, no script table.

90-second brief

How to read this

The low-code IDE still assumes someone will record, edit, and maintain a suite. Non-technical testers, PMs, and support still bounce off a script-shaped product. Web and mobile still live in different tools. I needed a Studios that removes the scripting layer, not another nicer editor for it.

  • Keep low-code as the only Studios

    Died

    Record-and-heal is a real product. It is still a barrier. We would keep selling to SDETs and losing the people who file the bugs.

  • Fully autonomous agent, no plan

    Died

    An agent that clicks before a human sees the plan is a liability in QA. Wrong path, wrong env, burned credentials. Fast demo. Unshippable for a release gate.

  • Chat-first, Plan mode, HITL

    Shipped

    Conversation is the authoring surface. The agent proposes a plan, waits, then drives a real browser or device. Pause and take over stay first-class. This is what shipped in 0.1.0.

Decision
Remove the scripting layer. Testing is a conversation with a plan you can approve. Web and mobile share one desktop app. Memory of the product is the category claim — not a one-off chat that forgets.
Result I will defend
Shipped on the skills-first stack: English in, structured bug report out. Chat, Plan mode, live view, OTP. Web and mobile in one macOS app. Weekly authors are not on a dashboard — I will not invent a seat count.
What I got wrong
I almost treated this as a skin on the low-code IDE. It is a different ICP. If the first screen is still a test table, the no-code bet is already lost. I also cannot pretend cloud device emulation is a device lab — Appetize is how we ship mobile without hardware. That constraint has to stay visible.
Who loses if this is wrong
The release manager, if an unreviewed agent path papers a defect. The tester, if chat cannot finish login. The SDET, if no-code floods suites nobody can debug.

Metric spine

North star
Time-to-first-executed-test
A PM or QA can describe a flow and see the agent driving a real browser or device without writing a locator. If we lose to “just ask ChatGPT and click along,” we lose the wedge.
Guardrail
Unreviewed action / silent drive
The agent must not burn a production login or a paid environment because Plan mode was skipped. Pause-before-act is a product requirement, not a preference.
Not instrumented
Weekly active authors who never open a script
That is the ICP proof. 0.1.0 does not have a published cohort. I will not invent one.
Not instrumented
Plan-accept rate vs takeover rate
Would show whether Plan mode is load-bearing or theatre. Not on a dashboard yet.
Not instrumented
Web vs mobile run mix
The category claim is both platforms in one tool. Split is not instrumented.

Remove the scripting layer. Conversation is the test. Plan, then drive — web and mobile in one app.

Low-code still asks someone to think in steps and locators. A fully unsupervised agent is a demo. The shippable product is chat-first authoring with a plan you can reject, a live view you can trust, and memory so the agent is testing this product, not a generic website. That is a different Studios from the recorder IDE — on purpose.

No-code agent loop
  1. Describe

    A test, a bug, or a plan in English.

  2. Plan

    Steps for approval. HITL can reject or edit.

  3. Drive + remember

    Live browser or device. Pause, takeover, persist context.