Musicfy Create a Song Guide: From First Idea to a Full AI Song
Jack RighteousCreate a Full Song With Musicfy
Start with an idea, your own lyrics or a musical direction. Use Simple when speed matters. Use Pro mode when you want direct control over lyrics, style, duration, tempo, key, time signature, writing behavior, reference audio and a trained voice.
Jump to a section
Choose by what you need
What is Musicfy Create a Song?
Inside Musicfy, open Create → Create a Song. It sits alongside other Musicfy tools such as voice conversion and instrumental creation.

Create a Song is designed to generate a complete song candidate. You can let Musicfy make many decisions for you or specify more of the musical direction yourself.
Comparing full-song generators? If your real decision is Musicfy or Suno, follow the hands-on Musicfy vs Suno comparison →. This guide stays focused on using Musicfy well.
You do not need music theory to begin. If you already understand BPM, key, time signature, arrangement or reference tracks, Pro mode gives you controls that let you use that knowledge.
Simple or Pro mode: which should you use?
| Simple | Pro mode |
|---|---|
| One main song description | Lyrics and musical style separated |
| Fastest idea-to-generation route | Better for controlled production |
| Track Length slider | Duration control |
| Instrumental option | Instrumental option |
| Optional trained voice where access supports Voice Targeting | Optional trained voice where access supports Voice Targeting |
| Best for exploration | Additional musical and reference controls where available |
Workflow 1: Create a song with Simple
1. Open Create a Song
Go to Create → Create a Song → Simple.
2. Describe the song you want
The Simple field asks for the song’s style, mood and subject. Treat it like a concise creative brief.
Warm uplifting pop song about finally believing in yourself, emotional male vocal, hopeful chorus, modern production.
Contemporary Afro-dancehall pop, warm bass, syncopated percussion, melodic male vocal, uplifting but grounded mood, restrained verses building into a wide memorable chorus.
3. Choose vocals or instrumental
Turn on Instrumental (no vocals) when you do not want a generated vocal. Musicfy’s current public Create a Song range lists instrumentals from 10 seconds up to 4 minutes. Treat older partner-frame wording about 10-second instrumental loops as launch-era information, not as the current instrumental ceiling.
4. Set the Track Length
This directly explained an issue in our early testing.

We asked for songs close to four minutes inside the written description, yet generations kept landing around 1:30. Reviewing the interface showed that Simple mode’s Track Length control itself was set to 1:30.
Practical lesson: when duration matters, set the dedicated duration control. Do not rely only on duration language inside a descriptive prompt.
5. Choose a trained voice if appropriate
Musicfy separates training a custom voice from using that trained voice to target a new full-song generation. Plan access can change, so use the current Musicfy plan screen as the authority for Voice Targeting availability. Inside Create a Song, the control has appeared as Sing in my voice in JR testing.
Need to create or validate a reusable voice first? Use the Musicfy Custom Voice Tutorial →
6. Generate—and listen like a creator
- Did Musicfy understand the genre?
- Does the vocal fit?
- Does the chorus feel distinct?
- Does the arrangement develop?
- Did the song reach the intended length?
- Are sections rushed or missing?
- What is worth keeping even if the whole result is not finished?
Workflow 2: Create a song with Pro mode

1. Lyrics — optional
Use the lyrics field when you want to supply your own words. JR’s August 2026 interface and partner-frame record described verbatim lyric handling in Pro mode. Because interface limits and behavior can change, confirm the current field guidance before a release depends on exact text handling.
Lyrics = what is sung.
Style = how the song should sound.
2. Style
The Style field prompts for genre, instruments and mood. In practice, that is only the starting point. A better way to use Style is to give Musicfy a compact sonic brief while letting dedicated controls handle settings such as duration, BPM, key, Literal/Creative behavior and reference strength when those controls are available in the current interface.
Style Prompt Engineering for Musicfy
JR field-tested One of the clearest lessons from our testing is that a dedicated control can matter more than descriptive text. Our repeated 90-second generations were not fixed by asking for a longer song in the prompt; the Track Length control was still set to 1:30. Treat that as a broader workflow rule: use each control for the job it is best designed to do.
Duration should control duration. Tempo should control tempo. Key should control key. Literal/Creative should control writing behavior. Reference/Follow should control how much the source influences the result when those controls are present. The Style field should not become a second copy of every setting on the page.
The JR five-part Musicfy Style formula
Define the musical neighborhood, preferably with a primary genre and useful crossover direction.
Mainstream reggae/dancehall crossover
Describe range, texture, authority, intimacy or emotional delivery when Musicfy is generating the singer.
Deep grounded male baritone, story-driven and restrained
Name the engine of the record: drums, bass, groove and the most important supporting instruments.
Punchy syncopated drums, controlled deep bass, warm guitar and organ
State the element that helps this song remain recognizable instead of becoming a generic genre exercise.
Expressive harmonica feature, call-and-response chorus
Describe the destination rather than making unverifiable promises such as “make this viral.”
Polished contemporary commercial mix with mainstream crossover appeal
Three levels of Style prompting
Reggae/dancehall, warm uplifting mood, male vocal, strong bass and drums.
Contemporary reggae/dancehall crossover, grounded male baritone, syncopated drums, warm deep bass, guitar and organ, uplifting anthemic chorus.
Contemporary reggae/dancehall crossover, deep grounded male baritone, punchy syncopated drums, controlled deep bass, warm guitar and organ, restrained story-driven verses opening into a wider gospel-soul chorus, expressive harmonica feature, polished mainstream production.
The progression is not “more words = better.” Each extra phrase should control a real creative decision.
What usually does not need to be repeated in Style
- Duration when the Duration control is already set.
- BPM when Tempo is explicitly selected.
- Key when Key is explicitly selected.
- “Be literal” when Literal is already selected.
- Repeated instructions to follow the source when Reference/Follow is already doing that job.
- Full lyrics or long story explanations.
- Detailed section commands that belong in lyrics/structure markup rather than sonic description.
Prompt conflict: avoid asking one field to solve everything
Style prompts can be too vague, but they can also become overloaded. Prefer selective specificity: a small number of decisions with clear jobs.
| If the problem is… | Change this first |
|---|---|
| Wrong overall sound | Style |
| Wrong duration | Duration / Track Length |
| Wrong tempo | Tempo |
| Wrong tonal center | Key |
| Source identity is being lost | Reference / Follow, then simplify competing Style instructions |
| Generated singer does not fit | Vocal wording in Style or Voice Targeting |
| One section is wrong | Use the narrowest current section-repair control available rather than rewriting the whole Style prompt |
Take the concept beyond Musicfy
3. Instrumental and voice
Choose instrumental-only generation when appropriate. If you want a trained custom voice to target a new full-song generation, use the current Voice Targeting option where your Musicfy access includes it.
Pro mode Advanced controls: more control, not more requirements

Duration
Musicfy’s current public site lists vocal songs from 45 seconds to 4 minutes. If length matters, use the dedicated duration control shown in your current interface rather than merely writing “four-minute song” in Style.
Tempo, time signature and key
JR testing has used dedicated musical controls in Pro mode. When they are present, set them only when repeatability or a specific musical target matters; Auto remains a legitimate creative choice.
Creative vs Literal
JR’s August 2026 interface exposed a Creative / Literal writing-style control. We treat its exact influence as a testable behavior rather than a guarantee of composition-level fidelity.
Reference Audio & Follow: What Does Musicfy Actually Preserve?
JR field record Musicfy has supported reference-audio style guidance in Create a Song. Exact upload limits and interface labels can change, so confirm the current reference-audio screen before relying on a specific file-size ceiling or control name.
What musical information are you giving Musicfy to learn from?
How much authority should that source have over the new generation?
Where do you want Musicfy to take the song sonically?
Reference adherence is not one score
When you say Musicfy “followed the reference,” separate the judgment into musical dimensions. A result can score well in one area and drift badly in another.
| Dimension | What to compare against the source |
|---|---|
| Melody | Did the recognizable melodic contour survive? |
| Vocal phrasing | Are words, entrances and rhythmic delivery similar? |
| Groove | Did the rhythmic feel and pocket survive? |
| Tempo feel | Does the perceived pulse remain similar? |
| Harmony | Is the tonal and chord movement recognizable? |
| Structure | Are verse, chorus, bridge and other sections still organized similarly? |
| Section timing | Do important sections arrive at similar points? |
| Instrument identity | Did signature instruments remain present? |
| Production character | How much did the sonic palette and energy change? |
| Emotional arc | Does the song build, release and resolve in a comparable way? |
| Duration | Did Musicfy preserve, compress or expand the runtime? |
Four different jobs for reference audio
Use the source to establish a musical neighborhood—energy, instrumentation, genre or production feel. Considerable melodic or structural divergence may be acceptable.
Keep the song recognizable while updating production. Melody, phrasing, structure, groove, signature instruments and emotional identity matter much more.
Keep the composition while rebuilding the recording. This is a stricter objective and we have not yet established that Musicfy performs it reliably enough to promise.
Use the source as a launch point for something intentionally different. In this workflow, divergence is part of the brief—not automatically a failure.
Before generating: decide what must survive
Do this before hearing the output. It prevents you from moving the goalposts after a generation surprises you.
Example: chorus melody, exact lyrics, central groove, signature harmonica, section order.
Example: drum sound, bass treatment, vocal production, stereo width, mix polish.
Example: intro texture, backing-vocal treatment, transitional effects.
The Follow lab: test influence instead of guessing
If you want to understand Follow, keep the same source, lyrics, Style and other settings as stable as practical, then change Follow deliberately.
| Pass | Follow | What to score |
|---|---|---|
| A | Low | Melody, rhythm, structure, instrumentation, duration, production change, drift |
| B | Medium | Same scorecard |
| C | High | Same scorecard |
| D | Maximum | Same scorecard |
Do not invent a sweet spot such as “75% is best” until testing supports it. The useful setting depends on what you are trying to preserve and what you are willing to change.
Reference × Style conflict test
This is one of the highest-priority questions still open. If the reference points one way and Style points another, which instruction wins—and for which parts of the song?
- Reference + minimal Style
- Reference + compatible Style
- Reference + moderately different Style
- Reference + strongly different Style
Run the sequence at one Follow level, then repeat at another. Compare whether melody, groove, arrangement, instrumentation and production respond differently as reference influence changes.
Source quality and source type are still open variables
We should not assume a dense mastered full mix behaves the same as a cleaner source. Future tests should compare a mastered song, premaster, instrumental, full vocal mix, isolated stem, short excerpt and complete track where rights and workflow allow.
Does Musicfy inherit the source duration?
Not yet verified In the Stone & Faith modernization, the source was roughly 3:15 while the generated version expanded substantially. That creates a specific question: does Musicfy treat source duration as information, or does the explicit Duration control carry more authority?
A useful test is to use the same reference while asking for the source length, a shorter version and a longer version, then compare section proportions, solos, chorus repetitions, intros and outros.
Signature element: presence is not the same as role
If the source contains a defining instrument or motif, score two things separately. First: did Musicfy preserve it? Second: did Musicfy preserve what it does in the song? A harmonica appearing somewhere is not the same as preserving its bridge feature, solo prominence, melodic identity or emotional function.
Source lineage: drift can become the next reference
Once an AI system changes the groove, arrangement or phrasing, that generated version becomes a new source if you feed it into another platform. The next model can inherit the first model’s decisions as if they were part of the song.
This is broader than Musicfy. It belongs to a professional AI production workflow: preserve your original source, label each generation, and know which version every later output came from.
What we know—and what remains open
| Status | Finding or question |
|---|---|
| JR field record | Reference-audio guidance has been used in Create a Song. |
| JR field-tested | Reference + Style can still produce a meaningful reinterpretation rather than a strict preservation. |
| JR field-tested | Modernization quality and source adherence can move in opposite directions. |
| JR field-tested | Drift introduced in one AI generation can carry into a second-generation workflow when that output becomes the next reference. |
| Not established | Higher Follow equals exact duplication. |
| Not verified | Maximum Follow preserves BPM, melody, section timing or arrangement exactly. |
| Unknown | Whether Follow affects every musical dimension equally. |
| Unknown | How source type and source quality change adherence. |
| Unknown | Whether short reference excerpts behave differently from complete songs. |
| High-priority test | Whether Style influence decreases predictably as Follow rises. |
Take the reference lesson into the wider JR system
If you cannot name what in a reference actually matters, the blocker may not be a Musicfy setting. Use the larger training to solve the underlying creative decision.
Lyrics vs Structure: Two Different Kinds of Adherence
JR field record JR’s August 2026 Musicfy record described Pro mode as preserving supplied lyric text. That answers only one question: did the system preserve the text? It does not, by itself, prove that Musicfy will preserve the song’s section order, number of repeats, section proportions, melody, phrasing, solo placement or transition timing.
Exact words, line order, repeated lines, parentheticals and ad-libs. This is the most literal meaning of “use my lyrics.”
Verse, chorus, bridge, solo and other section labels; their order; how many times each section occurs; and whether a section is omitted, duplicated or moved.
Where sections begin and end, how long they last, melodic phrasing, gaps, instrumental features, transition timing and the overall architecture of the finished recording.
Why this distinction matters
If every supplied word appears but the bridge is moved, the final chorus is doubled or the instrumental feature lands in the wrong place, calling the result “fully adherent” hides the actual failure. The same is true in reverse: a song may preserve verse → chorus → verse → chorus → bridge → chorus while paraphrasing a line or mishandling a parenthetical response.
For controlled AI music work, diagnose the layer that drifted before rewriting prompts or regenerating the entire song.
Section labels are instructions—not guarantees
Labels such as [Verse], [Chorus], [Bridge] and instrumental cues can communicate intended architecture. What we should not teach as settled fact is that every label is always interpreted the same way, that every bracketed cue will remain unsung, or that labels force exact section timing. Those behaviors require controlled testing.
| Layer | What to inspect | Example failure |
|---|---|---|
| Words | Exact lyric wording | A line is rewritten or omitted |
| Line order | Sequence inside each section | Two lines are reversed |
| Parentheticals | Responses/ad-libs | A response is ignored, merged into lead or overused |
| Section order | Verse / chorus / bridge sequence | Bridge arrives after the final chorus |
| Section count | Expected repeats | An extra chorus appears |
| Section proportion | Relative length of sections | A short bridge becomes a long new movement |
| Instrument cue | Placement and role | A solo is missing or appears somewhere else |
| Transition timing | Where sections enter/exit | Chorus arrives too early |
| Melody / phrasing | How words are performed | Correct words, different recognizable melody |
Stone & Faith: why one “adherence” score is not enough
Our modernization test gives us a practical example. The locked lyric sheet contains verses, repeated choruses, a bridge, a harmonica feature, a [long harmonica solo], a final chorus and parenthetical responses such as “hold tight,” “woah-oh-oh,” “yeah,” “bring it down,” “wear the crown” and “shake it now.”
Every lyric is correct, but Musicfy adds an extra chorus or moves the bridge.
The words and section order are correct, but the harmonica feature is absent, shortened or placed elsewhere.
A bracketed instruction such as [long harmonica solo] is sung aloud instead of being treated as a production cue.
The same sections occur in the right order, but the recognizable vocal phrasing or melody drifts significantly.
The controlled Lyrics vs Structure lab
To learn whether markup actually changes structural realization, keep the musical conditions stable and change only the way the lyric/section information is presented.
| Pass | Lyrics / markup condition | Keep fixed | Measure |
|---|---|---|---|
| A | Exact lyrics with minimal/no section labels | Reference, Style, duration, tempo, key, Follow and voice | Natural section inference |
| B | Exact lyrics with Verse / Chorus / Bridge labels | Everything else | Section order and repeats |
| C | Same lyrics + explicit instrumental/solo cues | Everything else | Cue placement and whether cues are sung |
| D | Same marked-up lyrics, regenerated | Everything else | Repeatability across generations |
For each pass, log omitted sections, duplicated sections, moved sections, extra repeats, parenthetical handling, solo placement, bridge realization, final-chorus count and overall runtime. One generation is evidence about that generation—not proof of a general rule.
A better scorecard for controlled songs
| Metric | Score | Question |
|---|---|---|
| Lyric wording | 0–10 | Were the supplied words preserved? |
| Line order | 0–10 | Did lines stay in the intended sequence? |
| Parenthetical / ad-lib handling | 0–10 | Were responses handled as intended? |
| Section order | 0–10 | Did sections occur in the intended sequence? |
| Section count / repeats | 0–10 | Were choruses, verses and other repeats counted correctly? |
| Section proportion | 0–10 | Did sections occupy roughly the intended amount of the song? |
| Instrumental cue placement | 0–10 | Did solos/features happen in the correct role and location? |
| Transition timing | 0–10 | Did section entrances and exits feel faithful? |
| Overall architecture | 0–10 | Does the finished song preserve the intended form? |
Fix the layer that actually failed
| What went wrong? | First place to intervene |
|---|---|
| Wrong or missing words | Lyrics field / lyric text |
| Wrong section order | Section markup, then Reference / Follow if preserving an existing composition |
| Extra or missing chorus | Structure markup + Duration; compare against source architecture |
| Solo missing or in wrong place | Instrument cue + Reference / Follow; use targeted repair if the rest is strong |
| Correct words but wrong melody / phrasing | Reference / Follow—not more lyric wording |
| Correct composition but one weak section | Use the narrowest available section repair before regenerating the whole song |
| Correct structure but wrong sonic character | Style |
Preservation priority for an existing-song modernization
When the objective is “same song, updated production,” lock the hierarchy before generating. A useful working order is:
This does not mean every project uses the same priorities. It means you should decide the hierarchy explicitly so that a shinier mix does not hide a composition-level miss.
What is confirmed, field-tested and still open
| Status | Finding or question |
|---|---|
| JR field record | August 2026 testing and partner material described verbatim supplied-lyric handling in Pro mode. |
| JR field-tested | Reference-heavy modernization can remain recognizable while arrangement and duration still drift. |
| Not established | Verbatim lyric handling guarantees exact song structure. |
| Not verified | Every bracketed section label or instrumental cue is interpreted consistently. |
| Not verified | Bracketed instrumental cues are never sung aloud. |
| High-priority test | Whether explicit section markup materially improves section order, repeat count and solo placement. |
| High-priority test | How repeatable the same marked-up lyric architecture is across multiple generations. |
Take the lesson beyond Musicfy
The distinction between words, structure and arrangement is platform-independent. It is part of building a reusable creative brief instead of treating every generation as a black box.
The JR professional workflow
- Know the objective. What are you actually trying to make?
- Establish a baseline. Generate before changing everything.
- Lock what already works.
- Change one meaningful variable.
- Compare the outputs.
- Document the result.
- Keep the best generation—not merely the newest one.
The generation-to-production bridge
This is the Musicfy implementation of the wider JR workflow: DEFINE → DIRECT → GENERATE → COMPARE → DIAGNOSE → REVISE → DOCUMENT → REUSE.
Musicfy Create Session Record v1
Use this after at least one real generation. The point is to make the next decision traceable, not to create paperwork for its own sake.
These are local working fields only. Nothing is submitted and entries do not persist after a reload. Keep the approved source/candidate and any permission or provenance evidence with the project. Use Print / Save PDF for a durable snapshot.
What we are field-testing next
- Duration consistency from 3:00–4:00
- Creative versus Literal behavior
- Follow strength at low, medium, high and maximum
- Reference × Style conflict
- Source type and source quality
- Duration inheritance from reference audio
- Lyric wording vs section-structure adherence
- Section markup and instrumental-cue interpretation
- Voice Targeting repeatability across trained voices and song contexts
- Genre-specific behavior
- Regeneration consistency
- Current section-repair behavior before stem separation
- Language and pronunciation behavior across selected languages
- Style prompt order, length, negative instructions and control redundancy
Last field-tested: August 31, 2026. Public product facts refreshed September 2026.
Create a Song checklist
- What is the song about?
- Do I already have lyrics?
- What should it sound like?
- How long should it be?
- Do BPM, key or meter matter?
- Do I need a trained voice?
- If yes, does my current plan include the needed Voice Targeting access?
- Do I need reference audio yet?
- If using a reference, what must survive?
- If structure matters, have I locked section order and repeat count?
- Was the duration correct?
- Were the lyrics handled properly?
- Were section order and repeats correct?
- Were instrumental cues placed correctly?
- Did the arrangement develop?
- Did it follow the intended style?
- Which reference dimensions were preserved?
- Is the failure textual, structural, arrangement-level or sonic?
- What single thing should I change next?
What comes after Create a Song?
A useful generation does not automatically need every other Musicfy feature. Choose the next branch according to the problem you hear.
Keep the useful song and use the narrowest section-repair option currently available before rebuilding everything.
Keep the useful song and work on the voice.
Use stems when you need to isolate a vocal, drums, bass or another layer.
Move into complete-track production, then run source, voice/model, plan and presentation checks.
Need the whole Musicfy system? Return to the Musicfy Creator Hub →
Make one song. Learn from the result.
The goal is not to prove that AI made a perfect song on the first attempt. The goal is to understand what you are trying to create, hear what Musicfy did with your decisions, and make the next decision more intentionally.
Musicfy Creator HubFree Creator AcademyCreate What You Love | Love What You Create.
Partner disclosure: Jack Righteous has an ongoing content and affiliate relationship with Musicfy. JR field-test and partner-frame observations are identified as such; public product facts were refreshed in September 2026. Product specifications can change, so verify current settings, plan access and model/voice permissions before a project depends on a specific limit or behavior.
↑ TopPut this into practice with Musicfy
Choose the workflow that fits your project. The Musicfy Creator Hub connects song creation, voice conversion, custom voices, instrumentals, stems and rights guidance.
Try Musicfy through my partner link
Affiliate disclosure: JackRighteous.com may earn a commission if you subscribe through this link. Choose a tool because it fits the work you want to make.