Case study 02 · AI product · independent build

One roster photo, a month of shifts

I designed and built an AI-enabled consumer product end to end, including a multimodal Gemini workflow that turns an unstructured roster photo into structured shift data the nurse confirms before it reaches her calendar.

30/30
cells read correctly when sending one row instead of the whole table
3–8 s
per scan, down from 30–50 s for the whole table
≈ ฿0.01
model cost per scan (~1.6k tokens, down from ~15k)
Pilot
iOS and Android preview builds; store launch pending

Role: product, UX, AI workflow design, Gemini integration and development — on my own. Built with: Expo (React Native) and TypeScript, Gemini, and AI-assisted development.

The user problem

Thai hospital nurses get their month as a ward roster: a table with every nurse on the ward down the side, every day across the top, and a short code in each cell — ช (morning), บ (afternoon), ด (night), a double like ช/บ, a day off, leave, plus the ward's own marks. It usually arrives as a photo or a screenshot in a LINE group.

To use a calendar app, a nurse finds her row and re-enters the month by hand — around 25 cells, every month, on a phone, often after a shift. It's repetitive, easy to get wrong by one column, and the moment where most people give up on a shift app.

The existing workflow

Before writing code, I studied a real nurse's month as she kept it in an existing shift app. Two things stood out. Doubles were the norm, not an edge case: 11 of 24 working days were double shifts, so a double had to be as fast to enter as a single shift — which is why WENUP models a double as two shift records rather than a flag. And her swaps with colleagues were recorded as terse free-text notes, a job being done with the wrong tool.

The roster photo itself was the other half of the workflow: the source of truth everyone already had, and that nobody's calendar could read.

Product hypothesis

If a nurse can turn the ward's roster photo into her month in under 90 seconds — and fix a wrong cell without leaving the review — she fills the month instead of abandoning the app. I set two bars before building the review screen: at least 90% cell accuracy on real sample rosters, and a full month imported in under 90 seconds.

Why AI was appropriate

I looked at the cheaper option first. Plain OCR returns words and their positions, not a table; rebuilding rows and columns from a skewed phone photo, with merged codes like ช/บ, superscripts and coloured cells, is the actual hard part. Google's on-device text recognition (ML Kit) doesn't support Thai script at all, so Android would have had no on-device path.

A vision-language model reads the table as a table. Given the image and the month, it returns structured data — the codes for each day — in a single call. That's where AI removes real effort: not as a chat feature, but as the step that converts an inconsistent, human-made document into data the app can use.

Why Gemini — and why the smallest model

Choosing a model was a cost decision as much as a technical one. WENUP sells to Thai nurses at ฿39 a month, and the plan for subscriptions assumes low willingness to pay, so every scan's model cost had to stay well below anything that would eat the margin on a Pro user who scans every month. My first estimate for a vision model, with Claude and Gemini on the shortlist, was ฿0.3–1 per scan: workable, but expensive for a ฿39 product if usage grew.

The constraint on the other side was quality in Thai. The roster codes are single Thai characters — ช, บ, ด — plus combinations like ช/บ and ward-specific marks, and a misread character becomes a wrong shift on someone's calendar. I needed a model that reads Thai script reliably, not one that's merely cheap.

Gemini's Flash-Lite tier met both: in row mode it read every Thai code correctly, with results identical to the larger Flash and Pro models, at about ฿0.01 a scan — 30 to 100 times below my estimate. So I shipped the smallest model that cleared the quality bar, and planned the larger Flash model as the fallback for any scan that fails the totals check. One more decision went with it: the key runs on Gemini's paid tier, because on the free tier Google may use submitted images to improve its products, and a ward roster is not something to hand over for that.

The AI extraction workflow

Scan pipeline

01

Photo or screenshot of the ward roster

02

She frames her own row; the app adds the day header above it

03

Gemini reads the strip and returns codes per day as JSON

04

The ward's codes map to her shift types; unknown codes are skipped

05

She reviews the month in calendar cells, then saves once

The first version sent the whole table. I tested it before designing the review screen, scoring each run against the totals the roster already prints for each nurse, so no hand labelling was needed.

Evaluation on real ward rosters (Sep 2026)
What was sentModelAccuracySpeed · cost
Whole table (26 rows × 30 days)Gemini Flash-Lite0 of 26 rows matched the totals5–25 s
Whole tableGemini Flash10 of 26 rows exact; ~11% of cells off, almost all a one-column slip; 3 of 9 calls timed out30–50 s · ~15k tokens
Her row + the day headerFlash-Lite, Flash, Pro30 of 30 cells correct on all three models3–8 s · ~1.6k tokens ≈ ฿0.01

That changed the product. Accuracy went from "needs a careful review" to "glance and confirm", latency from around 40 seconds to around 5, and cost dropped by roughly ten times. It also meant the other nurses' rows never leave the phone — the privacy decision and the accuracy decision turned out to be the same decision. Finding her row became the critical on-device step, so the scan screen is designed around framing one row.

A later bug made the same point. With neighbouring rows partly in view, the model sometimes transcribed a colleague's row instead — which surfaced as "my double shift wasn't detected" until a second scan. The prompt now names the strip's two sections explicitly, and the model reports how many rows it saw, so a crop that caught two rows is called out on the review screen rather than silently trusted.

Human-in-the-loop validation

AI output never edits a nurse's work schedule silently. The scan lands in a review step that looks exactly like her calendar, and nothing is saved until she confirms.

  • Unknown codes are left empty, never guessed. A wrong "leave" on a working day is worse than a blank she fills in.
  • The roster checks itself. The month as read is compared against the sheet's own per-nurse totals; every error seen in testing was a column slip that changes those totals, so a mismatch is the signal to look closer.
  • One correction model. She fixes a cell with the same tap-to-paint interaction she uses everywhere else — no separate editing mode to learn.
  • One undoable commit. The whole import saves as a single step that one undo reverses.
  • Each ward's codes are learned once. The first scan asks what the ward's codes mean; after that, they map automatically.

Key product decisions

  • Send one row, not the table. Chosen on evidence, for accuracy, speed, cost and privacy at once.
  • The smallest model that clears the Thai-quality bar. Gemini Flash-Lite at ≈ ฿0.01 a scan keeps a ฿39-a-month product's margins intact.
  • Skip rather than guess. Trust is the product: one confidently wrong shift costs more than a cell left blank.
  • Local-first by default. The calendar lives on the phone with no account needed, and backup is a file the user owns.
  • Roster images aren't kept. The image is used for the scan and not stored, and temporary copies on the phone are deleted when the scan screen closes.
  • Bad reads don't cost the user. A blurry photo or an outage gives the monthly scan back.

Monetization decisions

Pro is an App Store or Google Play subscription, or a one-time purchase: ฿39 a month, ฿349 a year (the default offer, with a 14-day trial), or ฿899 lifetime. Lifetime is deliberately priced as the anchor, about 2.6× annual, with room for a launch discount.

Roster scan sits in Pro because it is the one feature with a real cost per use. I moved widgets and shift alarms the other way, making them free: every competing shift app has them, and gating them would cost more trust than it earned.

Ads took two decisions. I first rejected programmatic ads for v1: in a market like Thailand they pay little, and an app whose best feature is a widget you don't have to open would have a standing incentive to make that widget worse. Later, I reversed that and accepted banner ads for free users, with firm placement rules: never on the scan, editing, sharing or payment flows, never on the alarm screen while it's ringing, and on the month screen only when the calendar keeps its full height. It's a judgment call about Thai users' willingness to pay, and the pilot will test it.

Built with AI-assisted development

WENUP was also an experiment in a different way of building products. Using AI-assisted development, I could move directly between product decisions, UX design and implementation, instead of handing requirements down a traditional product-development chain. I used different AI tools for different jobs, the way I'd staff a team.

Design. I designed the UI with Claude Design, working from written design briefs the way I'd brief a designer: the user, the constraints, the decisions already made, and screenshots of the working app attached each round. Eighteen numbered briefs took the product from the first month screen through roster scan, widgets, alarms, pricing and groups, so every design round started from what was actually built rather than from a blank canvas. For the illustrated theme system I moved into Figma, setting up the colour tokens as Figma variables with a mode for each theme in light and dark, so one token set drives every screen.

Build. The core logic — dates and overnight shifts, scan parsing and import planning, pricing entitlements, app-update rules — is covered by more than 40 automated test files, and the scan has its own evaluation scripts. The first commit was on 15 September 2026; three weeks later the repository had more than 330 commits.

Review — with a different model. I didn't let the model that helped write the code be the only one to check it. Branches went through a code review with Claude and a separate review with OpenAI's Codex, and the reviews before tester and store builds covered security as well as code. A second model has different blind spots, and the review rounds caught real issues before they reached testers — about 65 commits in the history are fixes from them.

AI changed how fast I could build. It didn't change what I was accountable for: the product decisions, the privacy model, the evaluation and what shipped were mine.

Results so far

WENUP is in pilot on iOS and Android preview builds, ahead of its App Store and Google Play launch, with a public landing page at wenup.app. I'm not yet reporting user numbers or conversion, because there aren't enough to mean anything.

What I can show is the evidence behind the AI feature: the row-only approach read 30 of 30 cells correctly on a real roster with every model tested, in a few seconds, for about one satang a scan — against a whole-table approach that slipped columns and timed out.

What I'd improve next

  • Photos of paper rosters. The evaluation so far uses screenshots and typed rosters. Paper photos need ground truth from a ward, and I expect cropping and deskewing to matter more there than the choice of model.
  • Finding her row automatically. On iOS, on-device text recognition can locate her name so she doesn't have to frame the row each month.
  • Measuring accuracy in the field. Track how often nurses correct a cell after a scan, to check the 90% target against real use rather than samples.

What I learned

The model is rarely the hard part. The hard part is deciding what to send it, what to do when it's unsure, and how the user stays in control.

The biggest improvement in this project came from changing the input, not upgrading the model — and it improved accuracy, cost and privacy at the same time. That's a product decision, and I could only make it quickly because I was close enough to the implementation to test it myself.