How ZWIP is built

Anyone can talk to an AI. Marking one is the hard part.

A general assistant will happily role-play a patient. It will not hold you to eight minutes, keep a marking scheme from you, or tell you that your history-taking is the weakest of your three domains. Here is what we built instead, and how we know it holds up.

15 free voice minutes when you register — no card required.

The station is the product, not the chat

A written scenario, not a prompt

Every station starts from an authored case: the presenting complaint, the patient’s history, their ideas, concerns and expectations, and how they behave if you rush them. You read the candidate briefing exactly as you would outside the exam room.

Examiner notes you never see

Each case carries hidden marking points and examiner guidance that are withheld from you and from the patient, then handed to the marker afterwards. You cannot be marked against a scheme you have already read.

Eight minutes, held strictly

The clock runs the way it runs in Manchester — reading time, then the consultation, then it stops. Timing is a skill the exam tests, so we make you practise under it rather than letting a conversation drift on until it feels finished.

Investigations released on request

Results exist behind the scenario and are handed over only when you actually ask for them. Ask well and you get them; forget, and you consult without them, exactly as you would on the day.

Everything above exists to reproduce one property of the real exam: the information is asymmetric. The patient knows things you have to earn. The examiner is holding a scheme you will never read. A conversation where everybody can see everything is a pleasant chat — it is not an OSCE station.

How the marking actually works

Three domains, marked separately

Data Gathering, Clinical Management and Interpersonal Skills are each marked in their own right, so a strong rapport cannot quietly rescue a weak history. You see where you lost the marks, not just that you lost them.

The total is derived, never asserted

Your overall score is computed from the domain marks rather than taken on trust, and score fields are protected from being edited after the fact. A grade that cannot be quietly rewritten is a grade worth arguing with.

A station you barely spoke in is not marked

If a consultation ends with almost nothing said, it is flagged as unassessable and kept out of your averages and pass counts instead of being scored as though it happened. Your trend line should reflect consultations you actually had.

One standard, every single time

Every consultation is marked by the same full rubric — there is no quicker, rougher path for a busy moment. If marking cannot complete, the station queues for another attempt rather than settling for less.

Marking we test, rather than trust

Most AI feedback tools ask you to take the grade on faith. We would rather show our working — so before our marking scheme is allowed anywhere near your results, it is run over a full set of 36 real consultations and measured. Two things matter in that test, and the rubric clears both.

It is consistent. Mark the same consultation twice and the domain scores move by an average of 0.10 out of 4 — a tenth of a mark. Across 72 repeated pairs, not one pass-or-fail decision changed. Practise the same station on Monday and Friday and the difference you see is you, not the marker having a different sort of day.

And it is steady across the board. The rubric is detailed enough that its marks barely shift with the machinery underneath it, which is what lets us keep improving the platform without your scores quietly drifting under your feet. A score from last month means the same thing as a score from today.

The whole calibration is re-run whenever anything about the marking changes, so an improvement in one place can never loosen the standard somewhere else.

Why the voice has to be fast

A consultation is a duet. You start a question, the patient begins answering before you finish, you hear worry in the pause and change direction. Put a two-second delay in the middle of that and the whole thing collapses into turn-taking over a walkie-talkie — you stop consulting and start dictating.

So the voice loop is built around latency rather than around features. Your speech is detected as you finish rather than after a fixed silence, the reply begins playing before it has finished being generated, and you can talk over the patient to interrupt them, exactly as you would if they were rambling in a real cubicle. None of that is visible on screen, which is rather the point: the measure of success is that you stop noticing there is a machine involved and start noticing you have three minutes left and no management plan.

So why not just use ChatGPT voice mode?

It is a fair question, and it deserves a fair answer. A general assistant is cheaper than we are, more flexible than we are, and genuinely good at things we do not attempt: explaining a guideline, drilling differentials, talking you through a management plan at midnight. If that is what you need, use it.

What it cannot do is examine you.

  • It has no marking scheme. Ask it how you did and you get an opinion, generated fresh, against no fixed standard — so the same consultation can be graded differently twice in a row, and nothing anchors either grade to PLAB 2.
  • It cannot keep secrets from you. Everything the marker knows, you have typed. Paste the examiner notes in and it can only mark the case in front of it, rather than the consultation you actually had.
  • It will not hold you to eight minutes. The conversation ends when you are satisfied, which is the one condition never available on exam day.
  • It has no station bank. You must invent the case, which means you already know the diagnosis, the hidden concern and the twist. You cannot surprise yourself.
  • It remembers nothing that matters. No domain scores, no trend across twenty consultations, no way to see that your interpersonal marks have been quietly sliding for a fortnight.
  • It wants you to feel good. A general assistant is trained to be agreeable. An examiner is not. That difference is not a tuning problem — it is the entire job.

None of this is a criticism of the technology. It is a statement about what a general-purpose tool is for. The marking scheme, the hidden notes, the clock, the station bank and the score history are the product; the conversation is just the part you can see.

What we don’t claim

  • We are not affiliated with the GMC. ZWIP is independent. Our stations and domains are modelled on published PLAB 2 material; nothing here is endorsed by the GMC.
  • A ZWIP score is not a prediction. It tells you how that consultation was marked against our rubric. It does not tell you what a real examiner will decide on the day.
  • No marker is perfect. Ours is measured for consistency rather than assumed to be right, and it holds to within a tenth of a mark between identical runs — but a genuinely borderline station is borderline, and one station is never the whole picture. Judge yourself on the trend across many.
  • An AI patient is not a person. It is a very good rehearsal partner that never tires and never has an off day. Practise with humans too, if you can.

More on how we write and review our content in our editorial policy.

Questions about how ZWIP works

Related reading

Judge it on your own results

Register free, run a real station, and see what a marked consultation looks like when the marker is not trying to make you feel good.

Start practising free

15 free voice minutes included — no card required