Esther MirzakhanyanProduct manager who vibe codes

Portfolio · 2026

Product manager who vibe codes.

I start from the person and the problem, decide what the product must never do, and then build the real thing with AI coding agents until it can be argued with. Three products, three regulated or sensitive categories, one method.

Esther MirzakhanyanSubscription products · experimentation · AI-built delivery
next slide previous F fullscreen
Nire telehealth home page on a phone
Nire · telehealth
IntimPath today screen
IntimPath · intimacy training
Lull today screen
Lull · baby sleep

How I work

PM who ships, with AI as the build capacity

Ten years across market research, product marketing and product management, the last four running subscription funnels on controlled experiments. In 2026 I started building the products myself with AI coding agents. The method stayed the same; the loop between an idea and a working screen shrank from weeks to hours.

  • Start from the person, not the backlog. Every decision traces to something a customer would say out loud. If I cannot name whose pain a feature removes, it does not get built.
  • Decide what the product must never do. In health-adjacent categories the "never" list is the product strategy: no invented proof, no claims the business cannot substantiate, no advice that needs a clinician.
  • Build the thing, not the ticket. I spec, build, review and QA with AI agents as engineering capacity, and I walk every changed flow in a real browser before it merges.
  • Distrust the output on purpose. Agents write confident comments about behaviour the code does not have. Rules become tests, claims are read against the code, and "green" means the case count, not the absence of red.
  • Leave it handoverable. A README you can run from, funnels and architecture as diagrams, and a checklist separating verified facts from assumptions.

Track record

Experimentation-driven growth, before the building started

Product manager on a high-traffic subscription platform since 2023: paid funnels, onboarding, paywalls, promotional placements and user segmentation, every release validated with a controlled experiment.

100+A/B tests run across web and mobile
+24%conversion and +15% revenue per user from subscription and acquisition flow work
+31%conversion and +17% revenue per user from redesigned onboarding and paywalls
35%experiment win rate, up from about 20%
+22%conversion from AI-powered personalization and recommendations
+28%campaign conversion after payment-history segmentation across billing, CRM and core configs
+19%revenue from localizing for Korea, Japan, Singapore and Taiwan
3products built end to end with AI coding agents in 2026, shown in this deck

Case study 1 · Direct-to-consumer telehealth

Nire: launching a regulated telehealth site fast and clean.

Weight management with compounded GLP-1s, hormone therapy, sexual health. A marketing site that had to be both compliant and convincing, and had to stay that way after launch.

The problem

A site that could sell, but could not launch

Paid acquisition was ready to start. The marketing site was not: the pre-audit build described compounded medications as "safe" and "evidence-based", promised "the same active ingredient" as branded drugs, showed a 4.7-star rating with 2,400 reviews that did not exist, and carried none of the disclosures the category requires.

The pipeline deployed straight to production. Every one of those claims was one merge away from being public under the brand's name.

Who was affectedPatients making a medical decision from marketing copy; the business carrying the legal exposure; the partner whose intake the site feeds
What was at stakeFTC action for unsubstantiated health claims, per-violation penalties for fake reviews, Florida weight-loss advertising rules, and a launch the growth team was waiting on
My roleProduct manager who vibe codes: I owned the problem, wrote the requirements, and shipped the code myself with AI coding agents

Why it mattered

In telehealth, trust is the conversion funnel

Compounded drugs are not FDA-approved, so claims of safety or efficacy are not marketing choices, they are legal exposure. Regulators had already moved against telehealth brands for exactly this pattern of claims and fake social proof. The same words that create risk also decide whether a patient trusts the brand enough to start an evaluation.

So the question was not "how do we soften the copy". It was: how does Nire launch on time with a site that is both compliant and convincing, and stays that way when nobody is looking?

  • Regulatory. Six marketing best-practice requirements from counsel, FTC substantiation rules, the fake-reviews rule, and Florida's weight-loss advertising rule.
  • Commercial. Paid social traffic was budgeted. Every week of delay was spend not deployed and a partner intake not fed.
  • Trust. Patients read the disclaimers. Real safety information and honest claims convert better than invented stars for a medical purchase.
  • Durability. A compliant disclaimer had already been deleted once in a routine commit. Without guardrails the fix would not survive the next release.

Goals and approach

Define "launch-ready" so everyone could check it, then audit, prioritize, fix once, lock it in

I turned the legal input into acceptance criteria that engineering, counsel and marketing could each verify: zero prohibited claims, every mandated disclosure on every weight-loss page, no social proof the business cannot substantiate, and guardrails in the pipeline. Checkout, intake and pharmacy vetting stayed explicitly out of scope so the launch did not expand.

  • Evidence-based audit. Every page checked against the six requirements. Each finding captured as a screenshot with route, rule and file, so the conversation with counsel was about facts.
  • Prioritize by exposure. P0 blocks any public deploy, P1 blocks marketing spend, P2 is hygiene. Four of six requirements failed; the P0 list was short enough to finish in days.
  • One tracker. 62 items across copy rewrites, additions, removals and process, each with exact replacement copy and status. The single source of truth for launch readiness.
  • Fix at the source. Copy lives in shared content fixtures, so one rewrite cleared the same claim on up to nine pages, and the design system kept every fix consistent.
  • Guardrails. The banned-claims list became a CI test; the review checklist lives next to the code.

Key decisions

Remove, replace, or add: the trade-offs

Compliance work is a series of product decisions, not a find-and-replace. These were the ones that shaped the launch.

  • Remove testimonials rather than replace them. There were no real reviews yet. A smaller home page with honest claims beat inventing social proof, and the section returns the day verified reviews exist.
  • Drop the trust seals until they are earned. LegitScript certification was not in place. A seal that links nowhere is a liability, not reassurance.
  • Add safety pages instead of footnotes. Patients were promised "important safety information" that rendered as plain text. Six real pages cost more than a link, and they are the most trust-building content on the site.
  • Rewrite claims to describe the process, not the outcome. "Prescribed and overseen by a licensed clinician" is true, specific, and still a strong reason to start.
  • Ship the disclaimer everywhere, not just where required. One footer component, one placement decision with the client, no per-page audit ever again.

Outcome

Launch-ready, and it stays that way

The site cleared every P0 and P1 item. The remaining tracker items are business inputs (pharmacy contact details, the substantiation file), tracked with owners and not blocking launch.

0prohibited claims on any page, enforced by the CI claims test on every merge
50 / 62tracker items landed; the rest are business inputs with owners
36routes on one design system of 62 components, including six new per-product safety pages
0engineering hours needed: specified, built and reviewed by the product manager with AI coding agents

What patients see now

Home

One hero message, the four treatment areas, clinician-led claims only. The rating widget and testimonials are gone; the footer carries the mandated disclosure.

Nire home page at desktop width
Homedesktop · 1440px

What patients see now

Treatment pages

A hub page plus one page per medication or treatment area, on the same templates with a different accent. Each weight-loss page carries the disclaimer, an "Individual results vary" line and a link to its safety page. Women's HRT stays "Coming soon" until the partner's intake supports it.

GLP-1 weight care hub page
Weight care hubGLP-1
Semaglutide treatment page
Semaglutideinjectable
Tirzepatide treatment page
Tirzepatideinjectable
Women's hormone therapy page
Women's HRTcoming soon
Testosterone therapy page
Testosterone therapymen
Erectile dysfunction page
Erectile dysfunctionoral

What patients see now

The pages that build trust

How it works, who we are, FAQ, contact, per-product safety information and the legal set in the same design shell. This is where a hesitant patient goes before starting an evaluation, so it got the same design attention as the treatment pages.

How it works page
How it works
Who we are page
Who we are
FAQ page
FAQ
Contact page
Contact
Semaglutide safety information page
Safety: semaglutide
Privacy policy page
Privacy policy

Before and after

Claim the process, not the outcome

Problem: "evidence-based" is an efficacy claim, and no clinical studies exist on compounded semaglutide. Decision: say what actually happens to the patient. A licensed clinician prescribes and oversees treatment, which is a real differentiator. "Individual results vary" was added under the call to action.

Before
Old hero: 'Evidence-based GLP-1 treatment'

GLP-1 hero. "Evidence-based GLP-1 treatment, reviewed by a licensed clinician and delivered to your door."

After
New hero: 'GLP-1 treatment prescribed and overseen by a licensed clinician'

Shipped. "GLP-1 treatment prescribed and overseen by a licensed clinician, delivered to your door." plus "Individual results vary."

Before and after

Earned trust in, borrowed trust out

Problem: the footer showed a LegitScript seal linking nowhere, a HIPAA badge and a pharmacy seal, none substantiated, while the disclaimer required on every weight-loss page was missing. Decision: remove the seals until the certifications exist, and add the disclaimer plus the Florida consumer bill of rights to the shared footer, so it appears on every page and cannot be forgotten on a new one.

Before
Old footer with unverified seals and no disclaimer

GLP-1 footer. Three trust seals, no disclaimer.

After
New footer disclaimer

Shipped. Full disclaimer text plus "Florida residents: see the Weight-Loss Consumer Bill of Rights."

Removed and added

Trade invented proof for real information

Problem: a 4.7-star rating, "2,400+ reviews", five placeholder quotes and a dead "View reviews" link, in a category where fake reviews carry per-violation penalties. Meanwhile "important safety information" was promised but rendered as plain text. Decision: remove the social proof entirely and invest the space in six per-product safety pages. Fewer stars, more trust, and a section that comes back the day verified reviews exist.

Removed
Old testimonial block with an invented rating

Home. Invented rating and review count with placeholder quotes.

Added
New semaglutide safety information page

Safety pages. What the medication is, who should not take it, side effects, and when to seek care.

Case study 2 · Consumer subscription · Sexual health

IntimPath: nobody books an appointment about this.

An intimacy-training app that takes what a good clinician would actually prescribe and turns it into a few private minutes a day on a phone. Nine programmes, 376 days of guided content, five languages.

The problem

Ordinary problems that almost nobody gets help with

Sexual problems are common and treatable. Yet it is embarrassing to say out loud, a clinic costs money and a waiting list, pills treat the symptom, and the internet contradicts itself. So people do nothing, for years.

The gap is not information; there is plenty. What is missing is a private, paced, personal plan: something that tells you what to do in the next ten minutes and makes you better at it over weeks.

4–15minutes a day, alone, on a phone
9programmes, one per situation
376days of guided content
5languages, because this is hard enough in your own

The user

Who buys this, and what actually hurts

Every decision in the product traces back to one of these sentences. A clinic means saying it to a stranger's face. Pills work while you take them and teach nothing. Free content is endless, contradictory and never sequenced. Generic fitness apps treat the pelvic floor as an afterthought and ignore the head entirely.

He says

"I finish too fast and I've started avoiding sex because of it."

"It works sometimes and not others, and I never know which tonight will be."

"I'm not going to my GP about this. I'm not taking pills forever either."

"I read that Kegels help. I have no idea whether I'm doing them right."

She says

"I love my partner and I just don't want it any more. I feel broken."

"I've never really known what works for my own body."

"Menopause changed everything and no one prepared me."

"Every article says something different and none of it tells me what to do today."

The insight

These are trainable problems, not fixed traits

Control, arousal and desire respond to practice the way strength does. The pelvic floor is a muscle that gets stronger with progressive load. Performance anxiety is a loop that breaks with attention training. Desire responds to stress, context and knowing your own body. All three need the same thing: short repeated practice, correctly paced, over weeks.

That is exactly what people cannot self-administer. They do not know the right tempo, when to progress, or what to do on a bad week. So the product's whole job is to hold the plan and set the pace, and to keep the effort small enough that someone actually comes back tomorrow.

Workout player with a ring that sets the tempo
Follow the ringIt sets the tempo and says when to contract and release.

The product

How each pain is answered

"I don't know if I'm doing it right"

The workout player

Pelvic-floor training fails because the muscle is invisible. A ring expands and contracts at the exact pace, a voice says contract and release, the phone can buzz the rhythm. Ten levels mapped day by day, so the load climbs without anyone deciding to make it harder.

"I start things like this and stop by week two"

The daily plan

The home screen shows today and nothing else: two short lessons and one or two sessions, four to fifteen minutes. No library to browse, no backlog to feel guilty about. Progress is one dial and a stage name, so a missed day is a pause, not a failure.

"It isn't only physical"

Mind and anxiety practices

Performance anxiety, a loud inner critic and stress are the real drivers for many people, so they are content, not a footnote: breathing, body scan, sensate focus, an inner-critic exercise and its counterpart that trains the voice that talks you down.

"Generic advice doesn't fit my body"

Nine programmes, assessed at both ends

Separate programmes for lasting longer, erection quality, pelvic strength, rebuilding desire, learning your own body, menopause, and a version written for queer women rather than a straight programme with the pronouns swapped. Each starts and ends with the same self-assessment, so the person sees what changed.

"I would never talk to anyone about this"

Private by construction

It runs in a browser: no app-store icon to explain, no appointment, no waiting room. Sensitive-data consent is asked for explicitly and recorded. For a category where shame is the main reason people never start, that is a product decision, not a technical one.

"Not in my language"

Five languages

Intimacy is hard to talk about in your first language and impossible in your fourth. Everything a customer reads or hears exists in English, German, French, Italian and Spanish, with the grammar written for the person reading it rather than machine-swapped.

In use

What the customer sees

Captured from the running application at phone size.

Today screen with two lessons and a workout
Just todayTwo lessons and a session. Nothing else to decide.
Lesson explaining the pelvic floor
Know what you're trainingTwo minutes on the muscle nobody explained to you.
Women's programme day
Her own planNot the men's programme with different pictures.
The same day rendered in German
In her languageAll of it, not just the buttons.
Library of programmes and practices
When they want morePractices for a bad week, off the daily path.

Growth

From a symptom to the right programme, and to the offer

Three quiz funnels of 52 to 60 screens turn a symptom into a programme, then signup, checkout and a fixed ladder of post-purchase offers. Three plans, one upgrade sold at two price points, prices editable without a deploy.

Opening screen of the quiz
First contactThe quiz that turns a symptom into the right programme.
Checkout page showing current and goal state
The promiseWhere you are now, where the plan takes you.
Upgrade screen
Going deeperAdvanced training for the people who stick.
Workout ring player
The habitThe session that brings people back tomorrow.

My part

What I owned, and what I shipped

I ran this as product manager and built it myself with AI coding agents: the decisions, the content, the funnel and offer copy, and the implementation behind them.

Decisions

  • Which programmes exist, how long each runs, and which stay hidden until they are genuinely finished
  • How the plan paces: what a day holds, when the level climbs, what a bad week looks like
  • Packaging and pricing: three plans, one upgrade at two price points
  • What the product may claim, in marketing and in the legal text

Delivery

  • 376 programme days written as structured content anyone can extend
  • Five languages, about 20,400 strings each, with agreement correct per audience
  • 616 illustrations, 82 audio tracks and 102 workout videos mapped to their slots
  • Three quiz funnels, signup, checkout and post-purchase offers

Keeping it honest

  • 335 automated tests where there were none, guarding the rules a content edit can break
  • A review and browser-QA pass on every branch before it reaches the CTO
  • Claims the product could not honour taken out of the legal text rather than left in
  • Documentation that lets the project change hands from one page

Case study 3 · Consumer subscription · Parenting

Lull: one honest companion for a baby's first year.

A sleep and care tracker where trust is the product: predictions as windows with a stated confidence, one table of truth for every number, and no invented social proof anywhere, enforced by a test.

The problem

It is 3 a.m. A parent is deciding, from a phone, about a person who cannot tell them anything.

  • When should the next nap be? The app says 14:20, as a fact. The baby disagrees. The parent stops trusting the app.
  • Is this a regression? Every rough week is labelled one, whether or not the evidence recognizes it, so the label means nothing.
  • What is normal at this age? Six sources, six answers, and the numbers on one site do not match the next.
  • What goes in their mouth first, what do they wear outside, does this rash matter? None of that lives in a sleep app, so the day is split across four tools and a group chat.
Today screen with the prediction dial
TodayThe next window, and how sure the app is.

The why

The parent does not need a more precise nap time. They need a source that never lies to them.

So trust was the product. Everything else, scope, pricing, architecture and copy, was derived from that one decision.

  • A wrong time costs trust in one afternoon. Sleep is noisy. Any product that pretends otherwise gets caught the first day the baby wakes early. Precision is not the job; honesty about uncertainty is.
  • Two numbers for one fact cost trust in one search. If the marketing page says one wake window and the app says another, the parent believes neither. In a health-adjacent product that is fatal.
  • Borrowed credibility costs trust for months. Ratings, testimonials and "% of parents" convert once. When a tired parent notices the numbers were invented, everything else the product says becomes suspect.

The bet

One honest companion for the whole first year, with sleep as the one domain that has a real engine

  • Whole day, one timeline. Sleep, feeds, nappies, solids and allergens, health and appointments, dressing for outside, growth, learning, and the parents themselves. The brief said "sleep tracker". The parent's day did not.
  • Free stays free. The tracker, every article and every age guide are free, forever. Premium is depth: the courses in full, the solids planner, long-range trends, export, more than one baby.
  • Three rules that never bend. Predictions are windows with a stated confidence, never a time. Anything about medicine or safe sleep is said plainly. No invented social proof anywhere, enforced by a test.
In, for the first versionThe tracker and the prediction engine. Two courses: sleep and starting solids. The solids planner and allergen ladder. Dressing by weather. Age-filtered guidance. A quiz-to-checkout funnel that runs end to end.
Out, on purposePrescribing a sleep-training method: the course explains them and the evidence, the app never asks you to do one. Assessing any symptom. A community feed. Anything that would need a clinician in the loop to be safe.
Why the line is thereEvery "out" is a place where being wrong would hurt a family. Those are not features to ship fast.

From why to what

A window, not a time. And an engine built to be honest about it.

  • Start from published norms. Age minus weeks premature picks one of 28 monthly bands. On day one the parent gets the population's answer, labelled as such. No empty state, because day one is exactly who needs it to work.
  • Learn this baby, slot by slot. Each wake window of the day is learned separately; the first window and the run-in to bedtime behave differently.
  • Never let one bad day move the estimate. Outliers are dropped, recent days weigh more, the norm keeps a hand on the result. Sick, travel, teething and vaccination days are shown but not learned from.
  • Change regime as the baby changes. Under six months wake windows drive; from nine months the clock drives; between, both. A nap drop is declared only after seven consecutive days of evidence.
  • Say which regressions are real. Four months is biological, eight to ten months developmental. Twelve months is a schedule problem, and the product says so.
Guidance screen naming the current wakeful phase
Guidance for this baby"Learning Mira's rhythm · day 2", then the phase the evidence recognizes.

From why to what

One table of truth. And the highest-stakes screen does the least.

The two trust rules from the "why", turned into architecture rather than left as policy.

  • One table of truth for age norms. The engine, the age guides on the public site and the course text all read the same per-month table. A page cannot quote a wake window the app disagrees with, because there is nowhere else for a number to come from.
  • The reaction logger never triages. When a food disagrees with a baby, the app offers the mild signs in the course's exact words, records the parent's choice, pauses the food, and stops. Severe signs go to the emergency number, never to a form. No severity score, no triage question. A test asserts the wording still matches the course.
  • Content is the product. The sleep course sits on an age-stage spine, because "mine is four months" is how a parent arrives. Each chapter is a ten-minute audio lesson for a parent who cannot read right now; the text is its summary; every module cites its source and ends in a quiz.

Growth

How a parent arrives, and why the first screen is a question

A three-minute quiz is the onboarding and the acquisition at once. It opens with the parent, branches on the baby's age, and feeds the answers straight into setup. The email is the account: no password, because a parent who abandons at "create a password" was never coming back. Three plan lengths, a fixed ladder of offers, and a decline path that returns to the exact step so a first purchase is never lost.

Quiz: who is filling this in
Who is filling this inThe parent first, then the baby.
Quiz: how old is your baby
Branch on ageNot born yet to two years.
Email screen
Email is the accountSign back in by link.
Checkout
CheckoutNow versus with a plan.
Payment declined screen
Decline pathBack to the exact step.

In use

The public site and the app, reading from the same table

Public courses, age guides and library pages quote the same per-month norms the engine uses. Inside the app: the timeline, the course, the guides and a plain "when to call your doctor".

Lull marketing home page
Public home
Courses page
Courses
Age guide page
Age guide
Course lessons screen
The sleep course
Log sleep screen
Log sleep
Four-month sleep change guide
What is normal
What to dress them in for outside
Dress for outside

Proof

What shipped in ten weeks

Counts read from the build, not typed in. A working beta: public site, funnel, checkout, app and API, walkable end to end on a laptop in a minute.

10kinds of entry on one timeline
28age bands of published sleep norms behind predictions, guides and courses
95course chapters across three tracks, every module cited and quizzed
46library articles, plus 16 month-by-month age guides
100first foods with age-banded preparation; 9 allergens on a ladder; 228 recipes
54quiz screens across two funnels, one engine
306automated tests, including the tone and no-social-proof rules
5locales routed; app strings translated

How it was built

Vibe coding, the disciplined version

  • Brief and rules first. What the product is, what it will never claim, how it speaks. Written before the first line of code, so the agents build against something.
  • Rules you can test. Tone, no social proof, wording that must match the course: each is a test, not a memo. A build that breaks a rule fails.
  • Counts, not adjectives. Anything the product claims about itself is computed from the build. If a number cannot be derived, it does not go on the page. Same for this deck.
  • I own what ships. Every diff reviewed, every changed flow walked in a real browser. The agents are fast, confident and sometimes wrong; my job is making wrong visible early.
What I would measure firstQuiz completion and email conversion. Share of babies whose prediction reaches "personalized" by day seven, which only happens if parents keep logging. Trial-to-paid. Course completion by module, to see where the audio earns its keep.
Gates before launchA clinician's review of the age table and the allergen method. A full visual redesign. The remaining audio lessons and photography. Translations beyond the app strings. Payments, CRM, support and legal.
StatusA beta, not a launch. I keep this as a checklist with a status per line, and I would rather show it than hide it.

What I bring

Product judgment, plus the hands to build the first version

  • Product. User research into decisions, scope and sequencing, pricing and packaging, funnels and conversion, retention mechanics, content strategy, launch readiness.
  • Experimentation. A/B testing programs, funnel and conversion optimization, user segmentation, success metrics and guardrails per experiment.
  • Sensitive and regulated categories. Writing about sex, health and the body without shame or clinical distance; consent, privacy and claim discipline.
  • Shipping with AI. Agent-driven implementation, review and QA gates, knowing where the output stops being trustworthy. LLM pipelines and LLM-based features in production.
  • Technical fluency. Node.js, React, SQL and migrations, git and merge-request flow, reading a diff and a stack trace.
  • Collaboration. Working to a strict reviewer, briefing design, CRM, billing and support, written handovers.

Experience

Oct 2023 — now
Product ManagerHigh-traffic subscription platform (under NDA). Paid funnels, onboarding, paywalls, segmentation, AI features.
2022 — 2023
Product Manager, YouScanB2B social listening. +18% free-to-paid on Dashboards, −19% churn after removing sales blockers.
2018 — 2022
Product Marketing Manager, DomenikLaunched brands, websites and campaigns; wrote requirements for design and engineering.
2015 — 2017
Market Research AnalystMarket, competitor and sales analysis.
Education
MBA, Wisconsin International UniversityPlus AI for Product Creation and A/B Testing at Projector Institute, GoPractice product management.

Contact

Looking for a PM who can build it?

I'm open to product roles where the problem is hard, the users deserve honesty, and a PM who can build the first version is an advantage. Happy to walk through any of these three products in detail, including the trackers, the tests and the parts that are not finished.

Esther Mirzakhanyan

Nire screenshots are from the launch build and the audit evidence set. IntimPath and Lull screens are from the running applications; Lull is pre-launch and prices are omitted on purpose.

1 / 1