A first recording is usually one unbroken take of ten minutes, read off a laptop propped against a mug, in a kitchen with the fridge running. Somewhere around minute nine comes a stumble and a muttered word, and the reader keeps going, because starting again means reading the whole thing again. On playback the fridge sits louder than the voice on every out-breath, and the muttering is still there at minute nine, where it stays until the file is deleted. What follows is the same job in the order the problems arrive, and none of it needs editing software, a microphone, or a quiet house.
What to gather before you press record
- The words, either a guided practice you can read off the screen or a script of your own already broken into short lines.
- A room with soft things in it: a made bed, a sofa, curtains, a rack of hanging coats. Sitting on the floor of a wardrobe is not a joke, and it is the best room most people own.
- A door you can close, and an agreement with whoever is on the other side of it.
- The noises you can switch off: the extractor fan, a laptop fan, the washing machine, and the phone itself set to Do Not Disturb so a notification does not land in the middle of a line.
- Water at room temperature, within reach, because a dry mouth is audible.
- Something to rest your forearm on, so the phone does not drift closer and further away while you read.
- Headphones for listening back, since a phone speaker hides the exact problems you want to hear.
Choosing or writing the words
The fastest honest start is a practice that already exists. Selfime ships 50 guided practices, six of them free, each already cut into cards with a pause written after every one, so the shape of the recording is decided before you open your mouth. Read one as written, or change the words on the card to sound like you. Changing “notice the weight of your hands” into a phrase you would say out loud is the single edit that makes a recording listenable on the ninetieth play.
If you are writing your own, the rule that matters most for recording is that every line has to be speakable in one easy breath, which for most people is somewhere between ten and twenty words. A line longer than one breath forces you to inhale in the middle of it, and the listener hears that breath as though something is about to happen. Writing the script well is a subject of its own and has its own post; for recording purposes, if the words fit one breath and point at something a person can feel or do in the room they are in, they will record.
You do not have to type any of it. The script screen can take dictation one segment at a time, using the speech recogniser on the phone, so a draft can be spoken, read back as text, tidied, and then recorded properly card by card. If you already have a practice you know well, skip the script altogether and record it straight onto a blank tape; with transcription switched on in settings the phone will write out each take afterwards, on the device and nowhere else.
The room, and where to hold the phone
A phone microphone is good enough. What ruins a recording is almost never the microphone and almost always the room, which is to say sound bouncing off flat walls and arriving back a fraction of a second late. Soft surfaces absorb that reflection, which is why a bedroom beats a bathroom and a wardrobe beats both, and why you want your back to the hardest wall rather than your face.
Then deal with the hum from a fridge, a heat pump, or a laptop fan, all of which sit in a low band that a phone records happily and that you stop hearing after half a minute in the room, which is why you will not notice it until playback. Listen to the room for a slow count of ten before you record anything, because whatever you can hear then, the recording will hear more of.
Hold the phone roughly a hand’s length from your mouth, about 20 centimetres, and slightly off to one side rather than straight in front of your lips. Speaking directly into a microphone drives the burst of air from p, b and t sounds at it, which lands on the recording as a thump rather than a consonant, and turning the phone twenty or thirty degrees off-axis removes most of them at no cost in level. Keep the distance steady, because moving from a hand’s length to half that between one card and the next is the most common reason two takes sound like two different people.
Pacing, and why the silence matters
The pacing already exists in the app, and it also tells you how fast to read. Read-Along holds each card on screen for as long as reading it aloud takes, with a floor so a short line still gets a moment, then holds the pause written into that card. The pace it reckons on is about 150 words a minute, unhurried speech rather than performance, so finishing a card well inside the time the app allows for it means you are going faster than the practice was written for.
The pauses are where the practice actually happens. A line like “let your shoulders drop away from your ears” takes three seconds to say and needs eight or ten for anything to come of it. Across a practice of a few minutes those pauses run from a few seconds to a good deal longer, and every one of them feels absurd while you are the person recording, watching a timer with nothing to do, and right the moment you are the person listening. The only fix for the doubt is to record one tape with the pauses as written, listen to it that evening, and notice that you wanted them longer.
The arithmetic is worth doing once, on twelve cards of roughly sixteen words, about six seconds of speech each, with eight seconds of silence after each one. That comes to about 77 seconds of voice and about 96 seconds of silence, a tape close to three minutes, more than half of it nobody saying anything. A closing card is given no pause at all, so a real tape lands a few seconds under that, but the balance holds, and it is not the balance most people expect when they sit down to record a three-minute practice.
Record one card at a time
Selfime records card by card rather than as one continuous read, which is what saves you from the ten-minute take, and each card can be replayed, redone, or kept before you move on, so a stumble costs one line instead of a session and you hear it immediately rather than at the end.
Recording in pieces also removes the reason people reach for editing software. There is nothing to cut, because a bad take is replaced by the next one, and nothing to stitch, because the pauses are held by the app rather than performed by you sitting silently in front of a running microphone. One long take makes every line a chance to ruin nine minutes of work, which is not a state of mind that produces a calm voice.
Read at the pace the card is written for, then stop. Resist the urge to leave your own silence at the end of a take, because the pause is added afterwards and any silence you record sits on top of it.
Ambience, and listening back
Ambience sits under the voice rather than beside it. Rain is free and the other soundscapes come with Lifetime, and the useful setting is quieter than it looks while you are choosing it, enough to cover a little room tone and give the pauses something to be, and no more. If you can still hear it clearly while the voice is speaking, turn it down.
Listen back on headphones, from the start, without stopping to fix things. What you are listening for is level, because one card recorded at a hand’s length and the next recorded from twice that will jump, and the jump is more distracting than either level on its own. Voice Balance evens playback levels between takes and keeps ambience behind the voice, which handles ordinary variation. It is not noise reduction, it will not remove the fridge, and it never rewrites the original recording, so a card badly out of line with the rest is one to redo rather than to balance.
Your recorded voice will sound wrong to you and to nobody else, because you normally hear it through the bones of your skull as well as through the air and the recording carries only the air part; that reaction fades with repetition faster than most people expect. The other thing worth saying is that a practice touching grief, panic, or a memory you did not plan to meet can land differently in your own voice than in a stranger’s. If recording or replaying something leaves you worse rather than steadier, stop the tape and talk to a clinician or a counsellor. Selfime is a recorder, not a medical device, and no guided practice is treatment.
The research nearby is thinner than the idea deserves. Kross and colleagues (2014) found that addressing yourself by name or as “you” rather than “I” reduced distress under social stress, and MacLeod and colleagues (2010) showed that words read aloud are remembered better than words read silently. Both are laboratory work with student samples and short tasks, neither studied recorded meditation, and neither shows that a practice in your own voice works better than one in someone else’s.
Where the tape ends up
The finished recording is encrypted in Selfime’s private app storage on the phone you recorded it on. There is no account, nothing is uploaded, and Selfime never receives a copy, which also means nobody at Selfime can get it back for you. Backup is an encrypted Recovery Kit you export yourself, opened with a 12-word recovery code, and that file can sit wherever you already keep things: iCloud Drive, Google Drive, Dropbox, a computer. Keep the code separate from both the phone and the kit, and make the kit the same day you record something you would be sorry to lose. Every privacy and recovery feature is in the free edition alongside its three recording slots, so the backup is not the part you pay for.
References
Kross, E., Bruehlman-Senecal, E., Park, J., Burson, A., Dougherty, A., Shablack, H., Bremner, R., Moser, J., & Ayduk, O. (2014). Self-talk as a regulatory mechanism: How you do it matters. Journal of Personality and Social Psychology, 106(2), 304-324. doi:10.1037/a0035173
MacLeod, C. M., Gopie, N., Hourihan, K. L., Neary, K. R., & Ozubko, J. D. (2010). The production effect: Delineation of a phenomenon. Journal of Experimental Psychology: Learning, Memory, and Cognition, 36(3), 671-685. doi:10.1037/a0018785