AI Coding
INTERNET CLAIMS. ACTUAL EXPERIMENTS.
Can Replit Really Build a Working App in Minutes With No Coding?
Replit says Agent can go from a single prompt to a working product in minutes and lets people build apps without knowing how to code. CASE 003 will test that claim across six apps with pre-written acceptance tests.
PRELIMINARY ASSESSMENT
INITIAL READ / BEFORE TESTING
Preliminary score. An initial reading of the claim, before the test.
THE INTERNET SAID WHAT? — CASE 003
APP IN MINUTES
The Claim
Replit currently markets Agent with a very direct promise:
Go from a single prompt to a working product in minutes.
It also tells prospective users they can build an app without knowing how to code.
That is an unusually clean Mythadis claim because we can define the requested application before the test, hand Agent one prompt, start a timer, and see what actually works.
The primary question for CASE 003 is:
Can a non-coder give Replit Agent one plain-English prompt and get a genuinely working app in minutes?
A secondary question is:
If the first build is imperfect, can the same non-coder get it over the line using only normal-language feedback and no manual code edits?
Preliminary BS Meter
4.4 / 10
This starts below the middle of the scale because the underlying capability is clearly real.
Replit Agent can create full applications from natural-language requirements, and Replit itself now evaluates vibe-coding systems with end-to-end functional tests rather than judging only whether code was generated.
The suspicious part is the compression in the marketing language:
- single prompt
- working product
- in minutes
- without knowing how to code
Each phrase may be reasonable on its own for some applications. Put together, they create a much broader expectation.
Replit’s own engineering team explicitly describes a functional-correctness gap: an AI coding agent can satisfy local coding constraints while the finished app still fails when a real user clicks around.
The working hypothesis is:
Simple applications will often work impressively fast, but success rate, required prompting, and build time will deteriorate as state, persistence, authentication, edge cases, and multi-step workflows are added.
Why We’re Suspicious
1. “Working” needs a definition
A page that renders is not necessarily a working application.
For this case, an app is working only if it passes the acceptance tests written before Replit sees the prompt.
Visual plausibility does not count as functional correctness.
2. “Single prompt” is different from “eventually works”
Modern coding agents are iterative by design.
If the first result contains broken flows and the user needs six rounds of:
- fix the login,
- fix the database,
- make the filter actually work,
- stop duplicate bookings,
then the tool may still be extremely useful—but that is a different claim from one prompt to a working product.
CASE 003 will therefore separate:
- first-prompt performance,
- no-code iterative performance.
3. “In minutes” may depend heavily on complexity
A landing page and a multi-user booking system are both “apps,” but they are not equivalent engineering tasks.
The test must include multiple levels of complexity rather than selecting one trivial demo.
4. Replit itself says end-to-end behavior is the hard part
Replit’s engineering team created ViBench specifically because conventional coding benchmarks can miss whether the application itself actually meets the product requirements.
Their own description says the relevant question for vibe coding is essentially whether the app loads, the core workflow works, and the result matches the request.
That is exactly what Mythadis should test.
5. No-code does not mean no judgment
A user may never type code and still need to understand:
- what is broken,
- what requirement was missed,
- how to describe the fix,
- when the result is actually safe to ship.
We should distinguish not writing code from not needing technical judgment.
What Would Change Our Minds
The Preliminary BS Meter should move down if:
- most apps pass nearly all pre-written acceptance tests after the initial prompt,
- simple and medium apps consistently become functional within a short elapsed time,
- complex apps remain surprisingly reliable,
- deployment works without manual code intervention,
- natural-language iteration fixes failures quickly when needed.
The meter should move up if:
- visually complete apps fail important workflows,
- the initial prompt regularly produces broken functionality,
- persistence/auth/state break under normal use,
- Agent repeatedly claims completion while acceptance tests fail,
- significant prompt iteration is required,
- manual code intervention becomes necessary,
- “minutes” becomes tens of minutes or hours for ordinary app specifications.
The Test
Build six applications from predefined product requirements.
Two simple, two medium, two more demanding.
Every application gets:
- one fixed plain-English prompt,
- a pre-written acceptance-test checklist,
- the same account/tier and Agent settings,
- no starter code,
- no manual source-code edits,
- timer started immediately before prompt submission.
App 1 — Quiz
Complexity: Simple
Build a five-question multiple-choice quiz.
Required behavior:
- one question at a time,
- four choices per question,
- cannot advance without selecting an answer,
- score calculated correctly,
- final result page,
- restart resets state.
Primary acceptance tests: 8
App 2 — Personal Habit Tracker
Complexity: Simple
Build a habit tracker where a user can create and track daily habits.
Required behavior:
- create habit,
- rename habit,
- delete habit,
- mark complete by date,
- completion survives page reload,
- basic streak calculation,
- empty state works.
Primary acceptance tests: 10
App 3 — Sales Dashboard
Complexity: Medium
Build a dashboard that accepts a supplied CSV of sales data.
Required behavior:
- import supplied CSV,
- calculate total revenue correctly,
- display sales by month,
- filter by product,
- filter by region,
- update displayed totals when filters change,
- handle malformed CSV gracefully,
- render at mobile width.
Primary acceptance tests: 12
App 4 — Event Signup
Complexity: Medium
Build an event-registration application.
Required behavior:
- list events,
- event capacity,
- attendee registration form,
- required-field validation,
- prevent duplicate signup using same email/event,
- capacity updates correctly,
- full event refuses additional registration,
- admin-style attendee view,
- persistence after reload.
Primary acceptance tests: 14
App 5 — Multi-User Expense Tracker
Complexity: Higher
Build an authenticated expense tracker.
Required behavior:
- account signup/login,
- logout,
- each user sees only their own expenses,
- create/edit/delete expense,
- categories,
- monthly total,
- category totals,
- persisted data,
- invalid form handling,
- second user cannot access first user’s records.
Primary acceptance tests: 16
App 6 — Appointment Booking
Complexity: Higher
Build an appointment-booking application for a small service business.
Required behavior:
- predefined available slots,
- customer booking form,
- booked slot becomes unavailable,
- prevent double booking,
- cancellation reopens slot,
- bookings persist,
- admin booking list,
- form validation,
- mobile usability,
- concurrent attempts at the same slot handled sensibly.
Primary acceptance tests: 16
Phase A — The Marketing Claim
For each application:
- Submit the complete predefined prompt once.
- Do not send follow-up prompts.
- Do not edit code.
- Allow Replit Agent to finish its autonomous work.
- Record elapsed time to its claimed completion.
- Deploy/publish if Agent’s normal workflow makes this available.
- Run the pre-written acceptance tests.
- Record every pass/failure.
Primary Phase A measurements
Per app:
- elapsed minutes,
- tests passed / total,
- percent acceptance-test pass rate,
- deployment success,
- obvious runtime errors,
- whether the core workflow works end-to-end.
Across all six:
- aggregate functional pass rate,
- pass rate by complexity tier,
- median build time,
- number of apps reaching the working threshold.
Working threshold
Pre-register:
An app counts as working from the initial prompt if:
- at least 90% of its acceptance tests pass, AND
- no failed test blocks the application’s primary workflow, AND
- it can be opened and used without manual code repair.
This is a Mythadis editorial test definition, not Replit’s own metric.
“In minutes” reporting
Do not create an arbitrary binary cutoff solely to manufacture a verdict.
Report actual elapsed time for every build.
For summary purposes also report how many working builds complete within:
- 5 minutes,
- 10 minutes,
- 20 minutes,
- 30 minutes.
That lets the audience decide what “in minutes” means without moving the goalposts after the test.
Phase B — No-Code Recovery
For every Phase A app that fails the working threshold:
Allow up to five follow-up prompts.
Rules:
- natural language only,
- no manual code edits,
- no direct database editing,
- tell Agent what behavior failed,
- do not tell it the implementation fix.
Example:
Booking the same slot twice is still possible. Fix that behavior and verify it works.
Not:
Add a unique composite index to table X and catch error Y.
Record:
- number of additional prompts,
- additional elapsed time,
- acceptance-test change,
- whether the app ultimately reaches working threshold.
This tests the broader without knowing how to code proposition separately from the strict single-prompt claim.
Phase C — Reality Check
For apps reaching the working threshold, perform a brief non-adversarial quality pass.
Check:
- refresh/reload,
- common invalid input,
- narrow/mobile viewport,
- second-user isolation where applicable,
- obvious error states,
- browser console errors,
- obvious exposed secrets if any external integration was introduced,
- whether the application still works after deployment.
This is not a full professional security audit.
Do not claim that passing Phase C makes an application production secure.
Reproducibility Rules
Before the experiment:
- Freeze all six prompts.
- Freeze all acceptance tests.
- Record Replit plan/tier.
- Record Agent version/mode when exposed.
- Record test date.
- Screen-record all six initial builds.
- Preserve Agent transcripts.
- Export/save application source where practical.
- Record elapsed times from the same start/stop rule.
- Do not alter acceptance criteria after seeing results.
For Phase B:
- preserve every follow-up prompt exactly,
- record which failed test motivated it.
What Happened
TEST NOT YET RUN
CASE 003 is currently in the experiment-design stage.
No performance result has been assigned.
The Verdict
PENDING
No final BS Meter will be assigned until the experiment has been run.
The Receipts
Primary Replit claims:
-
Replit Build: https://replit.com/build
Current language includes:
- “Go from a single prompt to a working product in minutes.”
- “Build an app without knowing how to code.”
-
Replit Agent: https://replit.com/products/agent
Replit describes Agent as building apps/sites through chat with no coding experience required.
-
Replit Agent 4: https://replit.com/blog/introducing-agent-4-built-for-creativity
Replit says Agent 4 is designed to help users ship production-ready software faster and presents “10X faster” as part of its launch positioning.
Technical/evaluation context:
-
Replit engineering — Closing the loop: Evaluating and improving Replit Agent at scale: https://replit.com/blog/evaluating-and-improving-agent-at-scale
Replit describes the functional-correctness problem in vibe coding and explains why it evaluates whether generated apps actually meet their product specs.
-
ViBench public benchmark: https://github.com/ViBench/vibench-public
The public benchmark contains PRD-based app-building tasks and natural-language functional test plans intended to evaluate whole-app behavior rather than code generation alone.
VIDEO: COMING
FINAL BS METER: PENDING
VERDICT: TEST IN PROGRESS
About the BS Meter: The BS Meter is a Mythadis editorial opinion based on the evidence available at publication. It is not a legal or regulatory finding and does not, by itself, allege fraud, criminal conduct or intentional deception. Editorial Standards
THE INVESTIGATION.
ON SCREEN.
External video platforms carry the broadcast. This case file keeps the investigation and its receipts together.
No video is attached to this investigation.