Vibe-to-Production · Framework

Score your app. Out of 100.

The same framework we score paid audits against, published in full and made tickable — all 49 checks, weighted, with the criticality of failing each one. Do it yourself, right here, before you speak to anybody including us. Then treat it as what it is: a baseline we tailor to where you are, never a script we run.

Score it yourself

Tick what is already true. Watch the number build.

Every check is here — nothing gated, nothing withheld to make the paid audit look better. Tick only what you have actually verified; a check you assume is fine is a check that fails. Nothing leaves your browser while you do this. How the weighting and the bands work sits behind How scoring works, next to your score.

0 of 49 ticked

Authentication & access control

0/20 pts

Secrets & configuration

0/10 pts

Data protection & recovery

0/10 pts

Automated testing

0/15 pts

Deployment & operations

0/15 pts

Performance & scale

0/10 pts

Payments & billing

0/10 pts

Code quality & maintainability

0/10 pts
Your production readiness
0/ 100
Not production-ready

Working demo, not a product. Putting real users or real data behind it carries risk you have not measured yet.

11 severe items still open
Diagnostic audit first — there are severe items open
The bands
  • Not production-ready029
  • Major gaps3049
  • Significant work remaining5069
  • Near-ready7084
  • Production-ready85100
Have this verified

Ticking sends nothing. Your score only reaches us if you fill in the form below. Findings come back in 72 hours, and we can sign an NDA before we see any code.

Have it verified

You scored it. Now let us try to break it.

Self-scoring is honest work, and it has one blind spot: you cannot adversarially test your own assumptions about who can reach what. Send us the score you just produced and an engineer will read it before anyone calls you.

Phone, with country code

WhatsApp works too if that is easier than a call.

We are not asking for your code. Read-only access gets arranged after we have spoken — and under an NDA first if you want one. Everything above is enough for us to tell you whether an audit is even the right next step.

Built from where you are

A baseline we tailor to you. Never a script we run.

This list is where an assessment starts, not what it is. Before anything gets scored, the checklist itself is rebuilt around your situation — categories that do not apply come out, obligations you carry go in, and criticalities shift with what your product actually holds. What to check first depends on which tool built the app; what to check hardest depends on everything below.

  • Your stage

    Pre-launch, a severe finding is a task on a list. With live users, the same finding is an incident and possibly a disclosure. The stage you are at changes which items block and which can wait.

  • Your stack & platform

    A design-led generator leaves a different gap than a full-stack one. What gets checked first — and how deep — follows from what built the application, not from a fixed order.

  • Your data sensitivity

    Health, financial or children's data pulls checks up a level and adds ones this list does not carry. An app holding public content is scored on a genuinely different bar.

  • Your obligations

    A compliance regime, an enterprise customer's security questionnaire, an investor's diligence checklist — whatever you have already promised gets layered onto the baseline before scoring starts.

What to do with it

From read access to shipped fixes.

Score it yourself honestly first — a check only passes if you have verified it, and our guide to triaging a production failure is free if something is already on fire. When you want it done adversarially, delivered and closed, this is the path our audit runs — each stage ends with something you hold, and every stage is priced before it starts.

  1. 01

    Get access

    Read-only · day one
    What happens

    You grant read access to the repository and infrastructure — under NDA if you want one. Nothing is changed, nothing is deployed. This is also where the checklist gets tailored: categories that do not apply to you come out, obligations you carry go in.

    What you walk away with

    A confirmed scope — exactly what will be assessed, what is excluded and why, before any clock starts.

  2. 02

    Audit

    Findings in 72 hours
    What happens

    Engineers assess your application against the tailored checklist adversarially — attempting the access that should be denied, reading the deployment configuration, testing the restore. Evidence, not opinion.

    What you walk away with

    A scored report where every finding carries its criticality, the evidence behind it, and the specific fix.

  3. 03

    Prioritise

    Ranked with you
    What happens

    Findings are ranked against your reality, with you in the room: severe items block launch, high items get dated, medium and important items get scheduled where they stop compounding. Your launch date and risk appetite set the order — not our template.

    What you walk away with

    A remediation plan with milestones that any engineering team could execute — including one that is not us.

  4. 04

    Ship the fixes

    Typically 2–6 weeks
    What happens

    Fixes land in criticality order, milestone by milestone — tests and the deployment pipeline go in early so every later fix ships safely. Each milestone is independently verifiable against the original finding.

    What you walk away with

    Closed findings, a re-scored application, and full handover — code, infrastructure and documentation are yours.

FAQ

Questions about the framework.

QWhat score is safe to launch on?
There is no universal number, and anyone who gives you one is selling something. A 70+ internal tool used by twelve colleagues is a very different risk from a 70+ application holding payment details for strangers. What matters is which specific checks failed and whether you have accepted those risks deliberately.
QIs this the same checklist you use in paid audits?
Yes — this is the framework, not a summary of it. What a paid audit adds is the evidence: engineers actually attempting the access that should be denied, reading the deployment configuration, and testing the restore. The list is the easy half to publish.
QWhy is authentication and access control weighted highest?
It carries 20 of 100 points because it is the only category where a single failure exposes every user at once. Everything else on this list degrades your product; this one can end it.
QHow are the criticality levels assigned?
Every check carries one of four levels — severe, high, medium or important — reflecting what a failure of that specific check costs, independent of its category's weight. Of the 49 checks, 11 are severe: the ones where a single failure exposes data, money or credentials and blocks a launch outright. The levels are our engineering judgement, published so you can argue with them.
QIs the checklist applied the same way to every application?
No — and that is deliberate. The published list is the baseline; before anything is scored we tailor it to your stage, your stack, the sensitivity of the data you hold and any obligations you have already taken on. Categories that do not apply come out, obligations you carry go in. A methodology applied identically to a pre-launch prototype and a live fintech product would be measuring the checklist, not the application.
QDo you need access to my code to score it?
Not to start, and not to have this conversation. Ticking the list here sends nothing anywhere — it stays in your browser. If you send us the score, all we get is the number, which items are open, and how to reach you. Read-only repository access is arranged after we have spoken, and we will sign an NDA before that happens if you want one.
QWhat happens after I send my score?
An engineer reads it before anyone calls you, so the first conversation starts from your actual gaps rather than a discovery script. Expect to hear from us within one working day. If an audit is the right next step, findings come back within 72 hours of getting read access — and if it is not the right next step, we will say so.
QCan I run this myself?
Yes, and you should before paying anyone. Most teams can honestly assess deployment, testing and code quality on their own. The categories where self-assessment tends to be optimistic are authorization and data recovery — because both look fine until they are tested adversarially.
QWhat if a category does not apply to us?
Payments is the common one — if you take no money, those points are excluded and the total rescales. We never score an application against something it does not do.
QHow long does a full assessment take?
We return findings within 72 hours of getting read access to the code. Running it yourself takes a focused day for most applications.
No-obligation diagnosis

Send us the repository. We'll tell you what's missing.

Fixed price, 72 hours to findings, and a report you can act on with or without us. Nothing is committed until you have read it.

NDA signed before you send anything · read-only access, revoked when the report lands · our copy deleted on delivery.