Back to work

Case study · Conceptual project

Designing an audio research notebook for collaborative student learning

Jess is a conceptual mobile app that helps university students turn audio into reusable study material: clip a useful moment from a podcast, audiobook, or uploaded text-to-speech file, get a transcript, save it to lab notes, and discuss it with a study partner around the exact timestamp.

Role
Self-directed — research, product framing, prioritisation, sketching, prototyping, informal testing
Type
Conceptual mobile app
Users
Research-heavy university students
Focus
Audio capture, transcripts, lab notes, collaboration, privacy
The MVP end to end — discover and play, clip and comment, record, save to lab notes, and share from a profile.

01 The problem

Students have good tools for studying from text. They can highlight a PDF, annotate an article, copy a quote, and share a reference that points to the exact line.

Audio gives them none of that. A useful idea sits inside a 40-minute episode, and unless the student stops to note a timestamp or switches to a separate notes app, it’s gone. Sharing is worse — a podcast link tells a study partner nothing about which two minutes mattered.

Studying from text

Highlight, annotate, quote, bookmark, share the exact line.

Studying from audio

None of that — the moment is buried, and links carry no context.

The question this kept coming back to: how might we help students capture, organise, and discuss useful moments from audio as easily as they already do with text?

02 Audience

The primary user is a research-heavy university student, using podcasts, recorded lectures, audiobooks, and uploaded readings while working on essays, assignments, and group study.

  • Goal — reuse specific points from audio in written work.
  • Pain — audio is hard to reference, scan, and share precisely.
  • Implication — capture, transcript, and retrieval have to be fast, and collaboration has to keep the timestamp attached.
Podcast user persona, Chris Farrelly, 24, a Trinity College masters student in cyber psychology who listens 2 to 3 hours a day, re-listens to grasp dense topics, and uses voice-memo apps to keep his thoughts in order. Goals include cataloguing information on the go and a more personalised, social listening experience; frustrations include juggling three apps to get the features he needs.
Research persona — Chris already cobbles capture together across voice memos and three different apps.

03 Research & insights

Research had two parts: a short survey of listeners about how they consume and study from audio (around two dozen responses, so I treat it as directional rather than robust), and desk and market research into how existing tools handle audio. I looked at how podcast apps, note-taking tools, read-it-later apps, and university learning tools break down once audio has to become study material.

Together, the survey and the desk research pointed to five conclusions, each of which shaped a feature.

What I concludedDesign response
Audio is easy to consume but hard to captureA fast clip action inside the player
Whole episodes are too broad to be useful referencesSave short timestamped clips, not whole files
Audio is hard to scan and search laterGenerate a transcript for each clip
Shared links carry no contextTimestamped comments tied to the exact moment
Study activity could feel exposing once social (a hypothesis I later tested)Private/public controls
Research snapshot: a listener survey covering age, listening frequency, why people listen, what they want from podcasts and audiobooks, when they listen and what apps are missing; alongside a desk and market synthesis of podcast industry trends, competitor apps, audio-clipping and social-audio tools, and industry commentary.
Research snapshot — the listener survey and desk/market synthesis behind the insights above.

04 Narrowing the MVP

The first version of Jess tried to be too much — discovery, social feeds, live broadcasting, premium courses. The most useful thing I did on this project was cut it back. I narrowed the MVP to one learning loop: capture a useful moment, give it context, and make it reusable. Everything that didn’t serve that was deferred.

DecisionWhat it covered
Kept (MVP)Player, clipping, transcripts, lab notes, study-buddy sharing, timestamped comments, private/public controls
SimplifiedTopic onboarding, discovery, text-to-speech upload
DeferredLive broadcasting, AI radio, premium courses, friends feed, profiles

05 Iterations

Four rounds of sketches show the narrowing in practice — the app moving from a broad podcast structure toward a focused clip-and-notes tool. Step through them; click a sketch to enlarge.

01 / 04
Hand-drawn sketch, iteration 1: three phone screens — search and discover, playlists, and a playlist detail — with a search / library / lab bottom navigation.
Iteration 01 — Podcast-app structure

06 The product loop

The MVP comes down to a single cycle that turns listening into reusable, shareable study material.

Listen Clip Transcribe Save Discuss Reuse a saved clip comes back when it’s needed
The MVP as one repeating loop.

07 Core flow — turning audio into reusable study material

The walkthrough follows one student from hearing something useful to reusing it later. Step through it below.

01 / 04
01 — Find the audio

08 Testing

I ran 3 informal usability sessions with friends who were students at the time. They weren’t formally recruited, so I treat them as directional rather than conclusive.

One signal stood out: a participant was uncomfortable that their study activity might be visible to others by default. That reinforced a decision I’d been weighing — to give students explicit private/public control over clips, notes, and playlists, rather than assuming a social-by-default model.

09 Trust & accessibility

I treated Jess as a multimodal study tool, not an audio-only app — students should be able to listen, read transcripts, search saved clips, and comment in text or audio.

  • Private by default, with clear sharing controls.
  • Source attribution on every clip.
  • Editable transcripts, so the text layer can be corrected.
  • Text alternatives for audio comments.
Risks I noted but didn’t solve ExpandCollapse

These would need real work before Jess could be more than a concept:

  • Copyright limits on clipping third-party audio.
  • Transcript accuracy, and how errors are surfaced and corrected.
  • Moderation of shared comments and playlists.
  • Proper support for Deaf and hard-of-hearing users beyond transcripts alone.

10 Outcome & reflection

Jess stayed a concept, so there are no adoption or usage numbers to report.

The real outcome was a decision: the strongest version of Jess wasn’t a broad social audio platform, it was a focused workflow for capturing, organising, and discussing spoken knowledge. Narrowing it was the part I’d point to as the actual UX work.

  • Can a student create their first clip without help?
  • Do they understand where it was saved?
  • Can they retrieve it later while writing?
  • Do they feel in control of who can see it?