Abstract
Asynchronous learning entered the applied music studio as an emergency substitute for the in-person lesson and has largely remained one. This paper argues that asynchronous material designed against established multimedia-learning principles can do work the live lesson cannot: extend the studio between lessons, make self-assessment part of the medium, and sustain study through the interruptions that characterise music learning in Southeast Asia. We set out the PACE framework — pacing, attention, contiguity, evaluation — a reorganisation of Mayer's principles of multimedia learning for one-to-one applied teaching, and report an exploratory enquiry conducted in the authors' own studios and through interviews with four musician-educators. The central proposition is that outcomes divide not by instrument and not by medium but by loop closure: material that returns student work to the teacher supports learning, while one-way delivery entrenches error. We state four hypotheses and a three-phase validation programme. This is a position paper; its claims are offered for testing rather than as findings.
The original had no abstract. An archive entry without one is hard to find and harder to cite.
Note the last sentence. Declaring the genre up front is the cheapest defence available against a reviewer who reads four interviews as an evidence base.
“General Session:” stripped — it dates the piece as minutes rather than scholarship.
PACE needs a collision check. The acronym is used in therapeutic parenting (Hughes) and in a speech-therapy protocol. Neither is in music or multimedia learning, but search before you commit. Mapping to your original four: instruction and visuals both resolve into attention and contiguity.
Section 01The default is not a design
Asynchronous teaching arrived in most private studios by necessity rather than by intent, and the shape it took then has largely stuck: a recording of what a lesson would have been, or a set of instructions that amount to prescribed homework.
The guiding question of this enquiry is whether asynchronous material can be something other than a substitute — whether, designed deliberately, it can carry pedagogical work that the live lesson is poorly placed to do.[1]
The literature offers a foundation but not an answer. Richard Mayer's twelve principles of multimedia learning describe how to design digital material so that a learner processes it without cognitive overload.[2] Varkey and colleagues applied those principles across K–12, college, graduate and continuing-medical-education cohorts.[3] Judith Bowman's The Music Professor Online supplied practical technique for post-pandemic music teaching in the tertiary institution.[4] What none of this addresses is the one-to-one applied studio — and, more pointedly for our purposes, none of it is drawn from Southeast Asia. The evidence base for online and asynchronous music pedagogy is overwhelmingly North American and European, tertiary, and institutional. The sector in which most music teaching in this region actually happens is private, independent, examination-driven, and almost entirely outside research attention.
This paper is an attempt to begin closing that gap. It is written by a vocalist and a pianist, each already drawing on the other's principal instrument in their own teaching.[5]
The origin story — the module on interconnection, meeting, mutual recognition — is compressed from 96 words to one clause plus endnote 5.
It was biography, not argument. Nothing in the paper's case depended on it.
This paragraph is the pioneering claim. The original never stated a gap, so the paper read as an application of Mayer rather than a first-of-kind contribution.
Naming the absence of regional, non-tertiary, private-sector evidence converts “we applied a framework” into “we established one for a context that had none”.
Section 02The Southeast Asian private studio
Four features of studio teaching in Singapore and Malaysia bear directly on how asynchronous material should be designed. They are not incidental context; each one changes a design decision.
Discontinuity is structural, not accidental
Study is routinely suspended during national examination years — PSLE, O-, N- and A-Levels, SPM — and a substantial proportion of students who suspend do not return. In Singapore, male students face a further documented two-year interruption for National Service. Where Western literature treats asynchronous provision as a convenience or an access measure, in this context it is first of all a retention mechanism: a way of holding a student in contact with an instrument through an interruption that would otherwise end their study.
Consumption is mobile-first
Personal media consumption in both countries is overwhelmingly mobile. This collides directly with the contiguity principle. A split screen carrying a score below and a hand close-up above is legible on a laptop and close to useless on a six-inch vertical display. Designs optimised for desktop viewing do not degrade gracefully to mobile; they invert the very split-attention load they were built to reduce. Re-specifying contiguity for small vertical screens is, in our view, the design problem most likely to produce a finding that travels back to the general multimedia-learning literature rather than merely borrowing from it.
Teaching is peripatetic and the practice environment is constrained
Many teachers travel between students' homes, absorbing travel time as unpaid cost; asynchronous batching converts that time into material-production time, which is an economic argument for adoption and not only a pedagogical one. At the student's end, high-density housing and noise tolerance limit practice to short and irregular windows. Material must therefore work in bursts — which is an argument for segmenting rather than a general preference for it.
Authority norms cut against autonomy claims
Student-led self-assessment presumes a willingness to evaluate one's own work against a teacher's model and to act on the difference without permission. That presumption is not culturally neutral. Where deference to the teacher is strong, autonomy may need to be constructed deliberately rather than assumed as a benefit of the medium. We raise this as a limitation on our own claims and as a hypothesis worth testing, not as a settled characterisation.
[AUTHORS: this section needs citation support — mobile penetration figures, attrition data around examination years, sector size and informality. Briefed as a separate literature session.]
The original contained no reference to Singapore, Malaysia or Southeast Asia. The stated ambition and the artefact were misaligned.
Four characteristics chosen from a longer list, on the principle that two or three developed properly beat ten listed. Examination discontinuity and National Service are the strongest because no Western literature has an equivalent.
Highest-value paragraph in the paper. Everywhere else you apply the literature; here you contradict it under regional conditions.
It also gives phase 3 of the validation roadmap a reason to exist.
Naming the deference problem yourselves is defensive writing in the good sense. A reviewer who spots an unqualified autonomy claim will use it; a reviewer who sees you raise it first will trust the rest more.
Section 03Why voice and piano: dual coding, not more senses
Singing is not recommended here because it adds a sensory channel. Within Mayer's dual-channel model — inherited from Paivio and Baddeley — learners process material through separate visual/pictorial and auditory/verbal systems, each of limited capacity, both requiring active processing. The benefit comes from distributing load across two channels, not from recruiting more senses. More is not better; balanced is better, and Mayer's own redundancy principle demonstrates that adding channels can impair learning outright.
Piano study routes almost everything through the visual and motor systems: notation, hand position, keyboard geography. Singing recruits the auditory channel for material that would otherwise queue behind all of it.
Its more important function, however, is diagnostic. Audiation, in Gordon's account, is comprehension rather than perception, and it is not directly observable.[6] Vocalisation externalises it and renders it assessable. A student who cannot sing a phrase has not yet internalised it, whatever the hands are doing. This is the argument for pairing voice with piano, and it is a stronger one than sensory accumulation.
The original claim — more senses engaged, more effectively retained — sits close to learning-styles folk theory, and Mayer's redundancy principle contradicts it. Citing Mayer in its support was the single most attackable sentence in the draft.
Rebuilt on dual coding plus audiation-as-diagnostic. The pedagogical case is now stronger, not weaker, for being narrower.
“Attempting the passage adds touch” is gone. Haptic and motor input are not channels in Mayer's model, so it could not be justified by the citation attached to it.
If you want the claim back, cite motor-learning literature for it separately.
Section 04The PACE framework
In the studio context we define asynchronous learning as learning that occurs outside real-time interaction, planned and supported by the instructor, with student autonomy and momentum as central motivators. Designed asynchrony is the practice of building that material against multimedia-learning principles rather than defaulting into it.
PACE reorganises Mayer's principles into four categories addressed to one-to-one applied teaching. It is a set of recommendations, not a checklist: each teacher will draw on different permutations according to teaching style and individual student need. The fourth column is the addition that matters most — naming the failure mode makes each row diagnosable rather than merely aspirational.
| Category | Mayer principle | Studio application | Failure mode |
|---|---|---|---|
| Pacing | Segmenting; learner control | Feedback delivered in short bar- or phrase-focused segments, each closing on one actionable point, so the student can replay, work, and return. | The single long take. Invites passive consumption rather than active practice. |
| Attention | Signalling; coherence | Bars highlighted and played in small sections; on-screen pointers at the moment of action; breath marks, melisma and improvisation cues aligned to the lyric. | Uniform emphasis. The learner must locate the point unaided, at cost. |
| Contiguity | Spatial and temporal contiguity; split attention | Score and hand close-up co-located; narration cued to the exact passage played; the student's submission aligned against the demonstration to reveal pitch and timing difference. | Score and demonstration separated in space or in time. See §2 on mobile. |
| Evaluation | Feedback; personalisation | Student submission and teacher response carried within the same medium; self-assessment against an aligned reference. | One-way delivery. Errors compound unobserved. See §5. |
Signalling, segmenting and contiguity were the three principles both authors leaned on most heavily in practice, and they form the spine of the examples above. Evaluation is the category the original model under-specified, and the interviews reported in the next section are the reason it now carries equal weight.
The table is the paper. The original described its central artefact and never printed it, which is the kind of omission a reviewer treats as fatal.
Now printed, captioned and numbered so it can be cited independently of the article.
The fourth column did not exist in your four categories. It converts recommendations into diagnostics — a teacher can look at a failing video and find the row.
It also carries the paper's argument forward: two of the four failure modes point directly at §5.
Instruction / visuals / pacing / feedback became pacing / attention / contiguity / evaluation so the set is memorable, acronymic and traceable to named Mayer principles.
Reverting is a find-and-replace if you prefer the originals — but you lose the name, and an unnamed framework cannot be adopted or cited.
Section 05Loop closure: what the interviews actually show
To test how far the model travelled beyond piano pedagogy, we interviewed four musician-educators — a saxophonist, a trumpeter, a vocalist and a bassoonist — each with roughly a decade of teaching or performance experience.[7] We asked what is hard to catch live but obvious on video, what logistical problems asynchronous work might solve, where it would not work, and what each would keep or redesign from the piano workflow.
Two accounts stand at opposite poles, and the distance between them is the finding.
The trumpeter had abandoned asynchronous lessons after students practised with incorrect technique and formed habits that then had to be undone. The vocalist described a workflow in which students recorded their singing and corrected it immediately, and reported it as complete rather than supplementary. Same medium; opposite outcomes.
The variable is not the instrument and not asynchrony. It is whether the loop closes. The trumpeter's students received one-way delivery: material went out, nothing came back, and error compounded invisibly. The vocalist's loop closed inside the medium itself. Note too what the trumpeter does now — recordings used as reminders of what was covered in person — which is precisely what this framework prescribes. Once the phases are separated, his case stops contradicting the model and starts confirming it.
| Phase | Suitability | Condition |
|---|---|---|
| Acquisition | Low | First encounter with new motor or technical content requires live diagnosis. Open-loop material is contraindicated. |
| Consolidation | High | Repetition and refinement of established material, provided the loop closes. |
| Review | High | Revisiting, self-assessment, and maintenance across interruption. |
The remaining two interviews sit consistently with this reading. The saxophonist valued revisiting modules and adjusting playback speed but cautioned that asynchronous material is inherently generic and may not locate an individual's root problem — support for review, not diagnosis. The bassoonist emphasised continuity beyond the lesson, tracking progress and preventing students from misremembering, and contributed the one concrete design recommendation to come out of the interviews: for wind instruments, keep the face on camera so that mouth shape and embouchure remain observable.
The draft reported this case and moved on, leaving an unreconciled contradiction with its own autonomy claims. It is the most valuable thing in your data.
Promoted from awkward exception to organising principle. “Open-loop asynchronous material is contraindicated for acquisition” is falsifiable, citable, and yours.
The acquisition / consolidation / review distinction is new. It resolves the contradiction structurally rather than by hedging the prose.
Endnote 7 now carries recruitment, sampling limitation, recording and consent. Keep it there rather than in the body — but it must exist somewhere, and in the draft it did not.
Section 06The somatic risk gradient
Across all four interviews the reported gains were varied while the reported limits converged on the body. We take that convergence seriously, but four interviews cannot support the conclusion that embodied content falls outside the medium.
Embodied content does not fall outside the asynchronous medium. It raises the cost of loop failure. Where errors are motor, self-reinforcing, and invisible to the learner, open-loop material does compounding damage rather than merely failing to help.
Read this way, somatic content predicts the severity of failure rather than the possibility of use, and the bassoonist's embouchure recommendation becomes a design response to a known risk rather than an unrelated tip. We offer it as a hypothesis for validation and not as a conclusion.
The draft's “live presence remains irreplaceable where the body meets the instrument” read as a hard boundary that your own data does not establish.
Recast as a graded risk. It now survives n=4, subordinates cleanly to loop closure, and gives the bassoonist's recommendation somewhere to sit.
Do not drop this. Four educators converging independently is a genuine signal; the problem was only the strength of the claim drawn from it.
Section 07Implementation and its trade-offs
A recurring aim of the session was to lower the perceived barrier for less technically confident teachers. Implementation requires no specialist software; the tools most studios already use are sufficient, and are set out in Appendix A.
The trade-offs are worth stating in the body, because they are pedagogical rather than technical. Recorded video is reusable and captures nuance, but introduces camera self-consciousness and measurable watch-through drop-off. Intentional instruction reduces cognitive load and builds independence, at the risk of over-engineering material that a spoken sentence would have handled. Asynchronous batching relieves make-up pressure and scales to larger studios, at the cost of upfront workload and blurred boundaries between work and home. A centralised coursework hub produces an organised student portfolio and gives parents visibility, but raises data considerations that should be handled on an opt-in basis, particularly where students are minors.
Zoom, Google Classroom, iMovie, Clipchamp and MuseScore moved out of the body. The tool list is the fastest-dating content in the paper and read as a workshop handout.
The four considerations stayed — those are argument, not tooling.
Deleted outright: “most teachers were already using at least two of these.” Unsupported, and the kind of claim a reviewer picks off easily.
Section 08Hypotheses and validation roadmap
This paper generates hypotheses; it does not test them. We state them explicitly so that they can be attacked.
- H1 · Loop closure. Asynchronous material incorporating a student-submission and teacher-response cycle produces better technical outcomes than one-way delivery of equivalent content.
- H2 · Phase. Open-loop material used for initial acquisition of motor-technical content produces higher rates of error entrenchment than the same material used for consolidation.
- H3 · Continuity. Asynchronous maintenance during national examination periods and National Service reduces attrition relative to suspension of lessons.
- H4 · Regional design. Contiguity designs optimised for desktop viewing lose effectiveness under mobile-first consumption, requiring re-specification of the split-attention principle for small vertical screens.
Phase 1 · Instrumentation and retrospective — 6 months
Formalise the interview protocol and expand the panel to 15–20 educators across Singapore and Malaysia, stratified by instrument family and studio type (independent, school-attached, chain). Retrospective attrition data from participating studios addresses H3 at low cost. Output: a validated instrument and a descriptive baseline.
Phase 2 · Multi-studio quasi-experimental — 12–18 months
Six to ten studios, students matched by grade level and years of study. Three conditions: open-loop material, closed-loop material, and lessons as usual. Measures pre-registered and mixed — examiner-blinded assessment of recorded performance, practice-log consistency, retention at twelve months, and student self-efficacy. Targets H1 and H2. Ethics review and parental consent are prerequisites, not formalities, given minors and recorded video.
Phase 3 · Regional design study — 12 months
Addresses H4. A/B comparison of contiguity layouts under mobile and desktop delivery, assessed by comprehension-and-execution checks. This is the phase most likely to return a finding to the general multimedia-learning literature rather than borrow from it.
What would disconfirm us
If H1 holds but H2 fails, the phase distinction collapses and the framework simplifies to a loop-closure requirement. If closed-loop material shows no advantage over lessons as usual on twelve-month retention, the autonomy claims in this paper should be withdrawn rather than qualified.
We invite studios in Singapore and Malaysia to participate in phase 1. [AUTHORS: contact address for participation]
A position paper without a roadmap is an opinion piece. This section is what converts the draft's assertions into a research programme.
Keep this paragraph even if you cut elsewhere for length. Stating what would falsify you is the fastest way to earn a reviewer's trust, and almost nobody does it.
The open call is the mechanism by which you become the centre of a regional research network rather than one contributor to it. It is a strategic move as much as a methodological one.
Section 09Conclusion
Designed asynchrony, organised through the PACE framework, treats material made between lessons as an object of design rather than a record of absence.
The paper generates three propositions for testing: that outcomes divide by loop closure rather than by instrument or medium; that open-loop material is contraindicated for the acquisition of motor-technical content; and that embodied content raises the cost of loop failure rather than placing content beyond the medium's reach.
To our knowledge this is the first framework of its kind developed from and for the private applied music studio in Southeast Asia — a sector that is large, independent, examination-driven, and effectively unstudied. The claims here rest on two studios and four interviews, and are stated at that strength deliberately.
What is required now is validation at scale, and we invite studios across Singapore and Malaysia to join it.
The original conclusion restated the body almost verbatim. This one does four jobs and nothing else: names the framework, states the hypotheses as hypotheses, makes the regional claim, and makes the ask.
Roughly 180 words against 130 — but none of it is repetition.
ApparatusEndnotes
[1] The model was developed and tested within the authors' own private studios between [dates] and refined through interviews with four musician-educators. No student outcome data was collected systematically during this period; the claims in §§3–6 rest on practitioner observation and are stated as such.
[2] Mayer's principles are presented in his own work as design guidance derived largely from short-form laboratory studies with undergraduate participants in expository science learning. Their transfer to one-to-one applied music teaching is an assumption of this paper, not a demonstrated result.
[3] Varkey et al. (2022) applied Mayer's principles across K–12, college, graduate and continuing-medical-education classrooms. None of the settings reviewed involved one-to-one instrumental or vocal instruction.
[4] Bowman (2022) addresses the tertiary institution rather than the private studio; her practical techniques transfer, her institutional assumptions largely do not.
[5] The authors met in a module on music in interconnection. Ahamed, a vocalist, turns frequently to the piano when teaching and composing; Chua, a pianist, incorporates vocalisation into her teaching. The collaboration followed from that overlap rather than from a prior research design.
[6] Gordon's formulation appears across his writing on Music Learning Theory; the phrasing here follows Learning Sequences in Music. The analogy between audiation and thought has been criticised for implying a stronger structural parallel between musical and linguistic cognition than the evidence supports. We adopt it as a working description of internal musical representation, not as a claim about shared cognitive architecture. [AUTHORS: verify page and edition]
[7] Participants were recruited by convenience from the authors' professional networks, which limits generalisability and probably biases the panel toward educators already sympathetic to digital tools. Interviews were [recorded / transcribed — confirm] and responses grouped by consensus between the two authors, with disagreements resolved by reference to the stated mechanism of the principle at issue rather than its typical application. [AUTHORS: consent and anonymisation arrangements; date range; any mid-study change to the protocol]
[8] The reorganisation of Mayer's twelve principles into the four PACE categories is a practitioner adaptation and has not been experimentally validated in this configuration. Category assignment was made by consensus between the authors.
The originals paraphrased sentences already in the body — they weren't doing endnote work.
Endnotes should carry four things the body can't hold: method detail, contested claims with the counter-position, provenance of decisions, and scope limits. Each one here now does at least one.
Note that several now concede something. That is deliberate — conceded limits are far cheaper than discovered ones.
ApparatusReferences
Bowman, J. (2022). The music professor online. [publisher, location — confirm]
Gordon, E. E. (2012). Learning sequences in music: A contemporary music learning theory. GIA Publications. [edition — confirm]
Mayer, R. E. (2021). Multimedia learning (3rd ed.). Cambridge University Press. [confirm edition used]
Varkey, T. C., et al. (2022). Asynchronous learning: A general review of best practices for the 21st century. [journal, volume(issue), pages — confirm]
[AUTHORS: regional sources for §2 to be added from the separate literature session — mobile penetration, examination-year attrition, private music education sector size.]
Four placeholders remain and the paper cannot be submitted with them. Also still missing: author designations and affiliations, and the permissions note for any figures or musical examples.
Nothing here is difficult — it is an hour of work that would otherwise be the first thing an archive editor sees.
Appendix AImplementation toolkit
Tool-current as of [date]. The framework is tool-agnostic; the table records what the authors used, not what is required.
| Function | Tool | Strength | Constraint |
|---|---|---|---|
| Scheduling · recording | Zoom | Already in use in most studios; recording built in | Storage limits on the free tier |
| Coursework hub | Google Classroom · Meet | Free; centralises communication; parent visibility; builds a student portfolio | Setup time; not music-specific; opt-in handling of student data required |
| Video editing | iMovie · Clipchamp | Free; low learning curve | Desktop-bound; limited for split-screen work |
| Annotated score | MuseScore | Free; supports the attention category directly | Notation-entry time |