← All work
The Princeton Review · Confidential, please don't redistribute · 12 min read
The Princeton Review LMS
Three years on one product: six exams, five course formats, and a million students a year. The whole arc, from the first audit to the numbers we reported to the board.
My role
UX/UI designer. Research through measurement.
Team
2 PMs, 6 engineers, curriculum, data, a second designer
Dates
Jan 2023 to May 2026
Scale
1M+ learners a year, 20+ programs
Contribution
User research
Moderated usability testing
Information architecture
Interaction design
Design system
Content design
Product metrics definition
Engineering handoff
Design QA
Tools used
Figma
FigJam
Jira
Confluence
Adobe Illustrator
Snowflake reports
Microsoft Teams
Lottie
Abstract
Between 2023 and 2026 I was the designer on a test prep platform used by more than a million students a year. Six exams and five course formats had drifted into thirty near-duplicate dashboards, and every student in a course received the same homework regardless of what their baseline test said they needed. I audited all thirty combinations, interviewed nineteen students across six exams, and rebuilt the platform on one shell and one design system. The finding that mattered most was not about layout. It was that our own vocabulary was telling students they were failing.
In short
Problem
Six exams and five course formats had drifted into thirty near-duplicate dashboards, each shipping its own vocabulary.
Users
Students studying alone between classes, and the instructors and advisors supporting them.
Discovery
An audit of all thirty configurations, separating real product gaps from years of drift.
Research
Nineteen moderated sessions across six exams against a written task rubric, plus two prototype directions.
Design
One shell, one progress model, one score-history model, and a design system behind 20+ programs.
Shipped
Ease of use held at 4.36/5 through a full navigation replacement; conversion 1.22% to 1.83% year over year.
Decision record
Goals
User: know what to study today and whether it is working. Business: lift free-to-paid conversion and keep students engaged past the midpoint of a course.
Constraints
One legacy platform shared by six exams and five course formats, release freezes during testing season, and 20+ live programs that could not break.
Iterations
Dashboard 2.0 to 3.0. Two prototype directions went to students; they preferred one's next-step hierarchy and the other's left navigation, so the shipped shell combines both.
Rejected
Keeping per-exam dashboards (it preserved the drift), and the most requested widget on the roadmap, a score prediction, because it implied a promise the business could not keep.
Edge cases
No baseline test yet, no test date picked, a test in progress versus taken versus not started, rest days in the plan, and Essentials students without live sessions.
Status
Live. The redesign is published on the Princeton Review platform.
What's on this page
1. The problem
2. How we worked
3. Discovery
4. Research
5. Synthesis
6. Wireframes
7. Design system
8. Validation
9. Build and handoff
10. Measurement
11. What came next
12. What I'd change
03
Discovery: product audit
Before designing anything I mapped every feature across every exam and every course format. Six exams, five formats, thirty combinations. I needed to know which differences were real product decisions and which were just drift.
Two things came out of it. Essentials courses were genuinely missing Advantage Sessions and flexible attendance, which is a roadmap problem, not a design one, so it went to the roadmap. Everything else was drift, and drift was mine to fix.
Live Online
In Person
Self-Paced
Essentials
Guaranteed
SAT
ACT
GRE
GMAT
LSAT
MCAT
Full feature set
Drifted: same capability, different nav or naming
Real gap
Not offered
Fig. 1. Feature audit across six exams and five course formats.
Three students, in their own terms
David, targeting 1500
Thirty minutes on weekdays, up to two hours at weekends. Full practice tests felt too long. He wanted one question per format instead of five repetitive ones, the real Desmos calculator built in, and equations shown under the prompt. He liked streaks because they felt like Duolingo. He said using a hint should not count as fully correct, and suggested partial credit based on how many hints you used.
Pedro, targeting 1400
Around twelve hours a week. He could not annotate or take notes on our content, and could not flag a lesson to come back to. He did not want the answer inside a hint. He wanted to finish the whole drill first, then see the correct answers and explanations together.
Eva, targeting 1250, hoping for 1400
Thirty minutes to an hour a day, chosen by energy level. Maths when tired, never reading. Highlighting and note taking kept her focused. She wanted to block out days she could not study and set her own study times. She tried to click into section progress on the study plan to reach the lesson behind it, and it was not clickable: a false affordance in the most-trafficked component on the page.
Renaming, after research
Not Ready
→Haven't started
Mastered
→Strong, 8 of 9 right
Projected score
→Your last test: 1390
Low yield
→Rarely on the test
LSAT students specifically told us they did not know what "low yield" meant.
Copy test, five students
"To begin, take Test 1 (Baseline Test)…"
"To get started, take Test 1 (Baseline Test)…"
A small thing, but it is the first sentence a student reads in the product. I polled the team and shipped the second one.
Two screens carry the decision the whole redesign turns on. The same student, the same week of work, in the two modes the shell supports.
Planner off is a sequence: the next class at the top, then lessons, then required homework, then extra practice. It answers what do I do next. Planner on is a calendar: the same items distributed across named days, with rest days the student chose and personalized drills that unlock after the next practice test. It answers when am I doing it. Interviews split cleanly on which of those two questions a student was asking, so rather than pick a winner I made it a toggle, and wrote the toggle copy to promise that switching loses nothing.
Fig. 3. Resource Library: practice tests with completion dates, scores, and view report actions.
Fig. 4. Module states. Four states, each carrying an icon and a label rather than color alone: done, in progress, where you are now, and not started.
Fig. 5. The practice test environment. It has to mirror the real exam exactly: the timer, the hide control, annotate, mark for review, and the question navigator at the bottom. Students told us a practice environment that differs from the real one is worse than no practice at all.
The library is not a swatch page. Every component below is specified with all of its states, because the states are where a learning platform actually lives: a class that is upcoming, live, attended, missed or rewatchable is five different pieces of copy and five different actions on the same card.
Three rules held it together. Status is never carried by color alone, so every state pairs a hue with an icon and a label. Activity type is carried by a single icon container whose fill identifies the kind of work, which is what let one card component serve lessons, drills, classes, books and personalized work. And every empty state is written, not defaulted, because a student who has taken no practice tests yet is the most fragile user on the platform.
Fig. 6. Foundations. Navigation, controls and form states, specified once and reused across six exams and five course formats.
Fig. 7. One card component, every activity type. Lesson, drill, live class, in-person class, book chapter and personalized work, each with the states the product actually produces.
Fig. 8. Progress and feedback. The four data states on the progress card matter most: no test taken, no target set, one of each, and both.
The component nobody asks you to spec
A student's practice test history is the only place they find out whether any of this is working. Ours was a table. I specced it as a chart with the target score drawn as a line, at three widths down to 375px, including what happens to the axis labels and the legend when there is no room for either.
The target line matters more than the trend. Score progression is not linear and a flat month is normal. Without the goal drawn on the same chart, a flat month looks like failure.
Fig. 9. Score history chart at three breakpoints.
11
Business impact
Four metrics, three commercial levers. Worth stating explicitly, because two of them read as satisfaction numbers and are really revenue and cost.
Revenue
Free to paid conversion is the only metric here that is directly money, and it moved from 1.22% to 1.83% year over year. On a base of more than a million annual learners, 0.61pp is a material lift, and the trial and practice-test journeys I designed sit directly in that funnel. I would not claim sole causation: pricing, marketing and curriculum all moved in the same period, so the honest framing is that design was one contributor to a measurable commercial result.
Retention
The disengagement definition changed when intervention was possible, from after a course ended to the halfway point. Every student caught at the midpoint is a completion the business keeps and a refund conversation it avoids. 35.6% of students were hitting that threshold, which is the size of the addressable problem, not a vanity figure.
Cost to deliver
One design system behind 20+ programs, replacing a parity spreadsheet and thirty near-duplicate dashboards. Six product teams stopped re-solving the same components, which is engineering time returned to roadmap work. It also meant a sitewide change, like the score-history chart, shipped once instead of thirty times.
Risk
Ease of use held at 4.36/5 through a complete navigation replacement. The commercial read on a flat satisfaction number is that a large redesign shipped without a support-cost spike or a churn event, which is the outcome a business actually buys when it funds a rebuild.