AI Music Creation: Step-by-Step Processes
Musicfy Create a Song Guide: From First Idea to a Full AI Song
Learn Musicfy Create a Song from first idea to controlled Pro generation. This practical guide covers Simple and Pro modes, lyrics, style, duration, tempo, key, reference audio, voice options, and a professional field-tested workflow.
Create a Full Song With Musicfy
Start with an idea, your own lyrics or a musical direction. Use Simple when speed matters. Use Pro when you want direct control over lyrics, style, duration, tempo, key, time signature, writing behavior, reference audio and a trained voice.
Choose your path
- Full songs can be generated from a prompt or from your own verbatim lyrics.
- Musicfy says Pro mode does not rewrite your supplied words.
- Vocal songs can run from 45 seconds to 4 minutes, with 48 kHz output.
- Musicfy lists 50+ supported languages, with 14 in live use during launch week.
- Reference-audio style guidance, Repaint and cover generation are listed as live features.
- Musicfy’s company-confirmed position is that downloads are not metered: no lifetime cap, no monthly allowance and no current plan to add one. This is not unlimited generation. Generation limits are separate.
What is Musicfy Create a Song?
Inside Musicfy, open Create → Create a Song. It sits alongside other Musicfy tools such as voice conversion and instrumental creation.

Create a Song is designed to generate a complete song candidate. You can let Musicfy make many decisions for you or specify more of the musical direction yourself.
Comparing full-song generators? If your real decision is Musicfy or Suno, follow the hands-on Musicfy vs Suno comparison →. This guide stays focused on using Musicfy well.
You do not need music theory to begin. If you already understand BPM, key, time signature, arrangement or reference tracks, Pro gives you controls that let you use that knowledge.
Simple or Pro: which should you use?
| Simple | Pro |
|---|---|
| One main song description | Lyrics and musical style separated |
| Fastest idea-to-generation route | Better for controlled production |
| Track Length slider | Duration: 45–240 seconds |
| Instrumental option | Instrumental option |
| Optional trained voice | Optional trained voice |
| Best for exploration | Tempo, key, time signature, Creative/Literal, reference audio |
Workflow 1: Create a song with Simple
1. Open Create a Song
Go to Create → Create a Song → Simple.
2. Describe the song you want
The Simple field asks for the song’s style, mood and subject. Treat it like a concise creative brief.
Warm uplifting pop song about finally believing in yourself, emotional male vocal, hopeful chorus, modern production.
Contemporary Afro-dancehall pop, warm bass, syncopated percussion, melodic male vocal, uplifting but grounded mood, restrained verses building into a wide memorable chorus.
3. Choose vocals or instrumental
Turn on Instrumental (no vocals) when you do not want a generated vocal. Musicfy’s partner publishing frame also lists 10-second instrumental loops as a current product spec.
4. Set the Track Length
This directly explained an issue in our early testing.

We asked for songs close to four minutes inside the written description, yet generations kept landing around 1:30. Reviewing the interface showed that Simple mode’s Track Length control itself was set to 1:30.
Practical lesson: when duration matters, set the dedicated duration control. Do not rely only on duration language inside a descriptive prompt.
5. Choose a trained voice if appropriate
Musicfy confirms that on Pro plans and up, a Create a Song generation can be sung in a voice you trained yourself. Musicfy identifies this capability as live since August 14, 2026. The interface exposes this through Sing in my voice.
Need to create or validate a reusable voice first? Use the Musicfy Custom Voice Tutorial →
6. Generate—and listen like a creator
- Did Musicfy understand the genre?
- Does the vocal fit?
- Does the chorus feel distinct?
- Does the arrangement develop?
- Did the song reach the intended length?
- Are sections rushed or missing?
- What is worth keeping even if the whole result is not finished?
Workflow 2: Create a song with Pro

1. Lyrics — optional
The current interface allows up to 4,096 characters and says, “Your words are used exactly as written.” Musicfy’s partner publishing frame independently confirms that Pro mode does not rewrite supplied lyrics.
Lyrics = what is sung.
Style = how the song should sound.
2. Style
The Style field prompts for genre, instruments and mood. In practice, that is only the starting point. A better way to use Style is to give Musicfy a compact sonic brief while letting dedicated controls handle settings such as duration, BPM, key, Literal/Creative behavior and reference strength.
Style Prompt Engineering for Musicfy
JR field-tested One of the clearest lessons from our testing is that a dedicated control can matter more than descriptive text. Our repeated 90-second generations were not fixed by asking for a longer song in the prompt; the Track Length control was still set to 1:30. Treat that as a broader workflow rule: use each control for the job it is best designed to do.
Duration should control duration. Tempo should control tempo. Key should control key. Literal/Creative should control writing behavior. Reference/Follow should control how much the source influences the result. The Style field should not become a second copy of every setting on the page.
The JR five-part Musicfy Style formula
Define the musical neighborhood, preferably with a primary genre and useful crossover direction.
Mainstream reggae/dancehall crossover
Describe range, texture, authority, intimacy or emotional delivery when Musicfy is generating the singer.
Deep grounded male baritone, story-driven and restrained
Name the engine of the record: drums, bass, groove and the most important supporting instruments.
Punchy syncopated drums, controlled deep bass, warm guitar and organ
State the element that helps this song remain recognizable instead of becoming a generic genre exercise.
Expressive harmonica feature, call-and-response chorus
Describe the destination rather than making unverifiable promises such as “make this viral.”
Polished contemporary commercial mix with mainstream crossover appeal
Three levels of Style prompting
Reggae/dancehall, warm uplifting mood, male vocal, strong bass and drums.
Contemporary reggae/dancehall crossover, grounded male baritone, syncopated drums, warm deep bass, guitar and organ, uplifting anthemic chorus.
Contemporary reggae/dancehall crossover, deep grounded male baritone, punchy syncopated drums, controlled deep bass, warm guitar and organ, restrained story-driven verses opening into a wider gospel-soul chorus, expressive harmonica feature, polished mainstream production.
The progression is not “more words = better.” Each extra phrase should control a real creative decision.
What usually does not need to be repeated in Style
- Duration when the Duration control is already set.
- BPM when Tempo is explicitly selected.
- Key when Key is explicitly selected.
- “Be literal” when Literal is already selected.
- Repeated instructions to follow the source when Reference/Follow is already doing that job.
- Full lyrics or long story explanations.
- Detailed section commands that belong in lyrics/structure markup rather than sonic description.
Prompt conflict: avoid asking one field to solve everything
Style prompts can be too vague, but they can also become overloaded. Prefer selective specificity: a small number of decisions with clear jobs.
| If the problem is… | Change this first |
|---|---|
| Wrong overall sound | Style |
| Wrong duration | Duration / Track Length |
| Wrong tempo | Tempo |
| Wrong tonal center | Key |
| Source identity is being lost | Reference / Follow, then simplify competing Style instructions |
| Generated singer does not fit | Vocal wording in Style or Voice Targeting |
| One section is wrong | Repaint / section repair rather than rewriting the whole Style prompt |
Take the concept beyond Musicfy
3. Instrumental and voice
Choose instrumental-only generation when appropriate, or use an available trained voice through Sing in my voice.
Pro Advanced controls: more control, not more requirements

Duration: 45–240 seconds
The current interface and Musicfy’s supplied product spec agree on 45 seconds to 4 minutes for vocal songs. If length matters, set it here rather than merely writing “four-minute song” in Style.
Tempo: 30–200 BPM
If you know the intended BPM and repeatability matters, set it. If you do not, Auto is a legitimate creative choice.
Time signature and key
Leave these on Auto unless meter or tonal center is a deliberate part of the test.
Creative vs Literal
Musicfy provides a Creative / Literal writing-style control. We are testing its exact influence rather than treating the label as a guarantee of composition-level fidelity.
Reference Audio & Follow: What Does Musicfy Actually Preserve?
Musicfy confirmed Musicfy supports reference-audio style guidance in Create a Song. The current Pro interface accepts MP3/WAV reference audio and displays a 100MB limit. What is not yet documented clearly enough to teach as settled behavior is which musical dimensions Follow preserves most strongly as you raise it.
What musical information are you giving Musicfy to learn from?
How much authority should that source have over the new generation?
Where do you want Musicfy to take the song sonically?
Reference adherence is not one score
When you say Musicfy “followed the reference,” separate the judgment into musical dimensions. A result can score well in one area and drift badly in another.
| Dimension | What to compare against the source |
|---|---|
| Melody | Did the recognizable melodic contour survive? |
| Vocal phrasing | Are words, entrances and rhythmic delivery similar? |
| Groove | Did the rhythmic feel and pocket survive? |
| Tempo feel | Does the perceived pulse remain similar? |
| Harmony | Is the tonal and chord movement recognizable? |
| Structure | Are verse, chorus, bridge and other sections still organized similarly? |
| Section timing | Do important sections arrive at similar points? |
| Instrument identity | Did signature instruments remain present? |
| Production character | How much did the sonic palette and energy change? |
| Emotional arc | Does the song build, release and resolve in a comparable way? |
| Duration | Did Musicfy preserve, compress or expand the runtime? |
Four different jobs for reference audio
Use the source to establish a musical neighborhood—energy, instrumentation, genre or production feel. Considerable melodic or structural divergence may be acceptable.
Keep the song recognizable while updating production. Melody, phrasing, structure, groove, signature instruments and emotional identity matter much more.
Keep the composition while rebuilding the recording. This is a stricter objective and we have not yet established that Musicfy performs it reliably enough to promise.
Use the source as a launch point for something intentionally different. In this workflow, divergence is part of the brief—not automatically a failure.
Before generating: decide what must survive
Do this before hearing the output. It prevents you from moving the goalposts after a generation surprises you.
Example: chorus melody, exact lyrics, central groove, signature harmonica, section order.
Example: drum sound, bass treatment, vocal production, stereo width, mix polish.
Example: intro texture, backing-vocal treatment, transitional effects.
The Follow lab: test influence instead of guessing
If you want to understand Follow, keep the same source, lyrics, Style and other settings as stable as practical, then change Follow deliberately.
| Pass | Follow | What to score |
|---|---|---|
| A | Low | Melody, rhythm, structure, instrumentation, duration, production change, drift |
| B | Medium | Same scorecard |
| C | High | Same scorecard |
| D | Maximum | Same scorecard |
Do not invent a sweet spot such as “75% is best” until testing supports it. The useful setting depends on what you are trying to preserve and what you are willing to change.
Reference × Style conflict test
This is one of the highest-priority questions still open. If the reference points one way and Style points another, which instruction wins—and for which parts of the song?
- Reference + minimal Style
- Reference + compatible Style
- Reference + moderately different Style
- Reference + strongly different Style
Run the sequence at one Follow level, then repeat at another. Compare whether melody, groove, arrangement, instrumentation and production respond differently as reference influence changes.
Source quality and source type are still open variables
We should not assume a dense mastered full mix behaves the same as a cleaner source. Future tests should compare a mastered song, premaster, instrumental, full vocal mix, isolated stem, short excerpt and complete track where rights and workflow allow.
Does Musicfy inherit the source duration?
Not yet verified In the Stone & Faith modernization, the source was roughly 3:15 while the generated version expanded substantially. That creates a specific question: does Musicfy treat source duration as information, or does the explicit Duration control carry more authority?
A useful test is to use the same reference while asking for the source length, a shorter version and a longer version, then compare section proportions, solos, chorus repetitions, intros and outros.
Signature element: presence is not the same as role
If the source contains a defining instrument or motif, score two things separately. First: did Musicfy preserve it? Second: did Musicfy preserve what it does in the song? A harmonica appearing somewhere is not the same as preserving its bridge feature, solo prominence, melodic identity or emotional function.
Source lineage: drift can become the next reference
Once an AI system changes the groove, arrangement or phrasing, that generated version becomes a new source if you feed it into another platform. The next model can inherit the first model’s decisions as if they were part of the song.
This is broader than Musicfy. It belongs to a professional AI production workflow: preserve your original source, label each generation, and know which version every later output came from.
What we know—and what remains open
| Status | Finding or question |
|---|---|
| Musicfy confirmed | Reference-audio guidance is a live Create a Song capability. |
| JR field-tested | Reference + Style can still produce a meaningful reinterpretation rather than a strict preservation. |
| JR field-tested | Modernization quality and source adherence can move in opposite directions. |
| JR field-tested | Drift introduced in one AI generation can carry into a second-generation workflow when that output becomes the next reference. |
| Not established | Higher Follow equals exact duplication. |
| Not verified | Maximum Follow preserves BPM, melody, section timing or arrangement exactly. |
| Unknown | Whether Follow affects every musical dimension equally. |
| Unknown | How source type and source quality change adherence. |
| Unknown | Whether short reference excerpts behave differently from complete songs. |
| High-priority test | Whether Style influence decreases predictably as Follow rises. |
Take the reference lesson into the wider JR system
If you cannot name what in a reference actually matters, the blocker may not be a Musicfy setting. Use the larger training to solve the underlying creative decision.
Lyrics vs Structure: Two Different Kinds of Adherence
Musicfy confirmed In Pro mode, Musicfy says supplied lyrics are used exactly as written and its partner publishing frame confirms that Pro does not rewrite the words you provide. That is important—but it answers only one question: did the system preserve the text? It does not, by itself, prove that Musicfy will preserve the song’s section order, number of repeats, section proportions, melody, phrasing, solo placement or transition timing.
Exact words, line order, repeated lines, parentheticals and ad-libs. This is the most literal meaning of “use my lyrics.”
Verse, chorus, bridge, solo and other section labels; their order; how many times each section occurs; and whether a section is omitted, duplicated or moved.
Where sections begin and end, how long they last, melodic phrasing, gaps, instrumental features, transition timing and the overall architecture of the finished recording.
Why this distinction matters
If every supplied word appears but the bridge is moved, the final chorus is doubled or the instrumental feature lands in the wrong place, calling the result “fully adherent” hides the actual failure. The same is true in reverse: a song may preserve verse → chorus → verse → chorus → bridge → chorus while paraphrasing a line or mishandling a parenthetical response.
For controlled AI music work, diagnose the layer that drifted before rewriting prompts or regenerating the entire song.
Section labels are instructions—not guarantees
Labels such as [Verse], [Chorus], [Bridge] and instrumental cues can communicate intended architecture. What we should not teach as settled fact is that every label is always interpreted the same way, that every bracketed cue will remain unsung, or that labels force exact section timing. Those behaviors require controlled testing.
| Layer | What to inspect | Example failure |
|---|---|---|
| Words | Exact lyric wording | A line is rewritten or omitted |
| Line order | Sequence inside each section | Two lines are reversed |
| Parentheticals | Responses/ad-libs | A response is ignored, merged into lead or overused |
| Section order | Verse / chorus / bridge sequence | Bridge arrives after the final chorus |
| Section count | Expected repeats | An extra chorus appears |
| Section proportion | Relative length of sections | A short bridge becomes a long new movement |
| Instrument cue | Placement and role | A solo is missing or appears somewhere else |
| Transition timing | Where sections enter/exit | Chorus arrives too early |
| Melody / phrasing | How words are performed | Correct words, different recognizable melody |
Stone & Faith: why one “adherence” score is not enough
Our modernization test gives us a practical example. The locked lyric sheet contains verses, repeated choruses, a bridge, a harmonica feature, a [long harmonica solo], a final chorus and parenthetical responses such as “hold tight,” “woah-oh-oh,” “yeah,” “bring it down,” “wear the crown” and “shake it now.”
Every lyric is correct, but Musicfy adds an extra chorus or moves the bridge.
The words and section order are correct, but the harmonica feature is absent, shortened or placed elsewhere.
A bracketed instruction such as [long harmonica solo] is sung aloud instead of being treated as a production cue.
The same sections occur in the right order, but the recognizable vocal phrasing or melody drifts significantly.
The controlled Lyrics vs Structure lab
To learn whether markup actually changes structural realization, keep the musical conditions stable and change only the way the lyric/section information is presented.
| Pass | Lyrics / markup condition | Keep fixed | Measure |
|---|---|---|---|
| A | Exact lyrics with minimal/no section labels | Reference, Style, duration, tempo, key, Follow and voice | Natural section inference |
| B | Exact lyrics with Verse / Chorus / Bridge labels | Everything else | Section order and repeats |
| C | Same lyrics + explicit instrumental/solo cues | Everything else | Cue placement and whether cues are sung |
| D | Same marked-up lyrics, regenerated | Everything else | Repeatability across generations |
For each pass, log omitted sections, duplicated sections, moved sections, extra repeats, parenthetical handling, solo placement, bridge realization, final-chorus count and overall runtime. One generation is evidence about that generation—not proof of a general rule.
A better scorecard for controlled songs
| Metric | Score | Question |
|---|---|---|
| Lyric wording | 0–10 | Were the supplied words preserved? |
| Line order | 0–10 | Did lines stay in the intended sequence? |
| Parenthetical / ad-lib handling | 0–10 | Were responses handled as intended? |
| Section order | 0–10 | Did sections occur in the intended sequence? |
| Section count / repeats | 0–10 | Were choruses, verses and other repeats counted correctly? |
| Section proportion | 0–10 | Did sections occupy roughly the intended amount of the song? |
| Instrumental cue placement | 0–10 | Did solos/features happen in the correct role and location? |
| Transition timing | 0–10 | Did section entrances and exits feel faithful? |
| Overall architecture | 0–10 | Does the finished song preserve the intended form? |
Fix the layer that actually failed
| What went wrong? | First place to intervene |
|---|---|
| Wrong or missing words | Lyrics field / lyric text |
| Wrong section order | Section markup, then Reference / Follow if preserving an existing composition |
| Extra or missing chorus | Structure markup + Duration; compare against source architecture |
| Solo missing or in wrong place | Instrument cue + Reference / Follow; repair the section if the rest is strong |
| Correct words but wrong melody / phrasing | Reference / Follow—not more lyric wording |
| Correct composition but one weak section | Repaint / section repair before regenerating the whole song |
| Correct structure but wrong sonic character | Style |
Preservation priority for an existing-song modernization
When the objective is “same song, updated production,” lock the hierarchy before generating. A useful working order is:
This does not mean every project uses the same priorities. It means you should decide the hierarchy explicitly so that a shinier mix does not hide a composition-level miss.
What is confirmed, field-tested and still open
| Status | Finding or question |
|---|---|
| Musicfy confirmed | Pro accepts supplied lyrics and says the words are used exactly as written. |
| Musicfy confirmed | The partner publishing frame says Pro does not rewrite supplied lyrics. |
| JR field-tested | Reference-heavy modernization can remain recognizable while arrangement and duration still drift. |
| Not established | Verbatim lyric handling guarantees exact song structure. |
| Not verified | Every bracketed section label or instrumental cue is interpreted consistently. |
| Not verified | Bracketed instrumental cues are never sung aloud. |
| High-priority test | Whether explicit section markup materially improves section order, repeat count and solo placement. |
| High-priority test | How repeatable the same marked-up lyric architecture is across multiple generations. |
Take the lesson beyond Musicfy
The distinction between words, structure and arrangement is platform-independent. It is part of building a reusable creative brief instead of treating every generation as a black box.
The JR professional workflow
- Know the objective. What are you actually trying to make?
- Establish a baseline. Generate before changing everything.
- Lock what already works.
- Change one meaningful variable.
- Compare the outputs.
- Document the result.
- Keep the best generation—not merely the newest one.
What we are field-testing next
- Duration consistency from 3:00–4:00
- Creative versus Literal behavior
- Follow strength at low, medium, high and maximum
- Reference × Style conflict
- Source type and source quality
- Duration inheritance from reference audio
- Lyric wording vs section-structure adherence
- Section markup and instrumental-cue interpretation
- Voice targeting and its relationship to duration
- Genre-specific behavior
- Regeneration consistency
- Repaint as a section-level repair before stem separation
- Language and pronunciation behavior across selected supported languages
- Style prompt order, length, negative instructions and control redundancy
Last field-tested: August 25, 2026.
Create a Song checklist
- What is the song about?
- Do I already have lyrics?
- What should it sound like?
- How long should it be?
- Do BPM, key or meter matter?
- Do I need a trained voice?
- Do I need reference audio yet?
- If using a reference, what must survive?
- If structure matters, have I locked section order and repeat count?
- Was the duration correct?
- Were the lyrics handled properly?
- Were section order and repeats correct?
- Were instrumental cues placed correctly?
- Did the arrangement develop?
- Did it follow the intended style?
- Which reference dimensions were preserved?
- Is the failure textual, structural, arrangement-level or sonic?
- What single thing should I change next?
What comes after Create a Song?
A useful generation does not automatically need every other Musicfy feature. Choose the next branch according to the problem you hear.
Try Repaint first when the rest of the song should remain intact. Musicfy describes Repaint as regenerating one section while keeping the rest.
Keep the useful song and work on the voice.
Use stems when you need to isolate a vocal, drums, bass or another layer.
Move into complete-track production, then run source, voice/model, plan and presentation checks.
Need the whole Musicfy system? Return to the Musicfy Creator Hub →
Make one song. Learn from the result.
The goal is not to prove that AI made a perfect song on the first attempt. The goal is to understand what you are trying to create, hear what Musicfy did with your decisions, and make the next decision more intentionally.
Musicfy Creator HubFree Creator AcademyCreate What You Love | Love What You Create.
Partner disclosure: Jack Righteous has an ongoing content and affiliate relationship with Musicfy. Musicfy-confirmed product statements on this page come from the August 2026 partner publishing frame supplied directly to Jack Righteous. Editorial workflow guidance and field-test findings remain independent. Product specifications can change; verify current settings before a project depends on a specific limit.
Develop the creative work
Turn the idea into a process you can repeat.
Find Your Sound connects song direction, revision, production decisions, packaging and release preparation.
Discussion