Finetuning.aiFinetuning.ai

GET /v1/songs/uploads/:id

An upload's status, its transcript and the tracks made from it

Fetch one song upload. Poll this after a Reimagine upload until status is transcribed, then read the lyrics we heard and the words we weren't sure of. It also lists every track made from the upload.

An upload stays readable after its 24 hours are up. By then deleted is true and audioUrl, transcript and lyrics are null, but generations still lists what you made from it.

Request

GET https://pub.finetuning.ai/v1/songs/uploads/:id

Headers

HeaderTypeRequiredDescription
X-API-KeystringYesYour API key

Path parameters

ParameterTypeDescription
idstringThe upload ID from POST /v1/songs/uploads

Example request

curl https://pub.finetuning.ai/v1/songs/uploads/22970e5b-c403-494c-b10d-846df6dde7c5 \
  -H "X-API-Key: ft_live_your_key_here"
async function waitForTranscript(uploadId: string) {
  for (;;) {
    const res = await fetch(`https://pub.finetuning.ai/v1/songs/uploads/${uploadId}`, {
      headers: { 'X-API-Key': process.env.FINETUNING_API_KEY! },
    })
    const { data: upload } = await res.json()
    if (upload.status !== 'transcribing') return upload
    await new Promise((r) => setTimeout(r, 5000))
  }
}

Response

200 OK
{
  "data": {
    "id": "22970e5b-c403-494c-b10d-846df6dde7c5",
    "feature": "reimagine",
    "name": "my-song",
    "format": "mp3",
    "status": "transcribed",
    "durationSeconds": 180,
    "originalDurationSeconds": 212.4,
    "window": { "startSeconds": 30, "lengthSeconds": 180 },
    "lyricsCap": 3240,
    "lyrics": "[verse]\nLate night on the empty street\nNeon humming to the beat\n\n[chorus]\nHold on, hold on\nWe were never gone",
    "transcript": {
      "language": "en",
      "sections": [
        {
          "label": "Verse",
          "lines": [
            { "words": [{ "w": "Late", "low": false }, { "w": "night", "low": false }, { "w": "on", "low": false }, { "w": "the", "low": false }, { "w": "empty", "low": false }, { "w": "street", "low": false }] },
            { "words": [{ "w": "Neon", "low": true }, { "w": "humming", "low": false }, { "w": "to", "low": false }, { "w": "the", "low": false }, { "w": "beat", "low": false }] }
          ]
        },
        {
          "label": "Chorus",
          "lines": [
            { "words": [{ "w": "Hold", "low": false }, { "w": "on,", "low": false }, { "w": "hold", "low": false }, { "w": "on", "low": false }] },
            { "words": [{ "w": "We", "low": false }, { "w": "were", "low": false }, { "w": "never", "low": false }, { "w": "gone", "low": false }] }
          ]
        }
      ],
      "lowConfidenceCount": 1,
      "wordCount": 19,
      "noVocals": false,
      "confidenceSource": "heuristic"
    },
    "lowConfidenceWords": [
      { "section": 0, "line": 1, "word": 0, "text": "Neon" }
    ],
    "fingerprintStatus": "skipped",
    "errorMessage": null,
    "pricing": { "transcriptionPaid": true, "transcriptionCreditApplied": true, "nextReimagineCredits": 2, "retranscribeCredits": 1 },
    "audioUrl": "https://dev.finetuning.ai/media/uploads/22970e5b-…?exp=1790866400&sig=…",
    "createdAt": "2026-09-30T13:53:20.151Z",
    "expiresAt": "2026-10-01T13:53:20.151Z",
    "deleted": false,
    "generations": [
      {
        "id": "959ef62f-45c1-4cb6-a4db-fa1da227f5fe",
        "type": "reimagine",
        "status": "completed",
        "title": "my-song (Custom)",
        "style": "Custom",
        "audioUrl": "https://media.finetuning.ai/audio/965f8c8d-….mp3",
        "errorMessage": null,
        "credits": 1,
        "usedTranscriptionCredit": true,
        "refunded": false,
        "createdAt": "2026-09-30 13:53:44"
      }
    ]
  }
}

The upload object

FieldTypeDescription
idstringUpload id — pass it as uploadId
featurestringreimagine or remove_vocals — what it was uploaded for
namestringFrom the file name, without the extension
formatstringmp3 or wav
statusstringSee below
durationSecondsnumberLength of what we kept (at most 180)
originalDurationSecondsnumberLength of the file you sent
windowobjectstartSeconds and lengthSeconds of the part we kept
lyricsCapnumberMost lyric characters Reimagine accepts for this upload
lyricsstring | nullThe transcript as lyrics with section tags — what Reimagine sings when you don't send lyrics. After a Reimagine it holds the lyrics that were last used
transcriptobject | nullSections → lines → words, each word with low: true when we weren't sure of it. noVocals: true means we heard no singing
lowConfidenceWordsarrayEvery uncertain word with its position (section, line, word are zero-based indexes into transcript.sections)
fingerprintStatusstringskipped, clear, match or error — the catalogue check, where enabled
errorMessagestring | nullWhy the last step failed, when it did
pricingobjectWhat the next steps cost. nextReimagineCredits: 1 while this upload's paid transcription credit is unused, else 2. retranscribeCredits: what /transcribe would cost now (0 for the free retry). transcriptionPaid / transcriptionCreditApplied: whether a transcription was paid for, and whether a Reimagine has used its credit
audioUrlstring | nullA signed link to the audio we kept, valid until an hour after expiresAt. null once deleted
createdAt / expiresAtstringUploaded at / deleted at (24 hours later)
deletedbooleantrue once the audio and transcript have been deleted
generationsarrayTracks made from this upload, newest first (up to 20). Only on this endpoint; the list and create responses return []. Each has credits (what that render charged) and usedTranscriptionCredit (true for the first render, which cost 1 on top of the transcription)

Status

StatusMeaning
readyStored. A Remove vocals upload stays here; it doesn't need the lyrics
transcribingWe're listening for the lyrics (usually 10–90 seconds)
transcribedLyrics ready in lyrics and transcript
no_vocalsWe didn't hear any singing. Send your own lyrics to Reimagine it
failedWe couldn't listen to it. Try again, or send your own lyrics
blockedRefused by the catalogue check and deleted

How sure are the "low" flags? Our speech recogniser doesn't report per-word confidence yet, so today (confidenceSource: "heuristic") we flag the words that look wrong: garbled characters, letters from a script the rest of the song isn't in, and the stock phrases speech recognition invents over silence. When per-word confidence is available the same fields carry it (confidenceSource: "words" or "segments"). Treat the flags as "worth checking", not as a complete list.

Errors

CodeStatusDescription
NOT_FOUND404No upload with that id on your account

On this page