Dictating a task is faster than typing one: “Create a bug in ENG for the login redirect, high priority.” This guide builds an n8n workflow that takes a recorded voice note, transcribes it with OpenAI, pulls out the task details as structured data, and creates the task in Jira or Asana, then replies to whoever sent the audio with what it created.
Updated October 2026: I rebuilt this guide. The 2025 version told you to resample every recording to 16 kHz mono with ffmpeg through n8n’s Execute Command node, to call “the hosted Whisper API”, and to use the legacy Move Binary Data node. Since then: n8n 2.0 turned the Execute Command node off by default and it doesn’t exist on n8n Cloud; OpenAI’s transcription API takes common formats directly (so the conversion step isn’t needed); OpenAI now recommends newer transcription models over Whisper, which it calls legacy; and n8n’s built-in OpenAI node still transcribes with whisper-1 only (I checked the node’s code in n8n 2.41.4), so this version calls the API directly to choose a model. I checked the workflow’s parameters against the n8n 2.41.4 node packages and tested the validation script, but I did not run it against live OpenAI, Jira or Asana accounts for this update, so expect to adjust IDs and names.
How the Pipeline Works
- A client (a phone shortcut, a small app, a script) sends an audio file to an n8n Webhook protected by a secret header.
- n8n sends the audio to OpenAI’s transcription API and gets text back.
- An Information Extractor node asks a language model to fill a fixed schema from the text: platform, project, summary, description, priority and assignee. It returns data, not prose.
- A Code node validates the result, fills safe defaults and stops with a clear error when something essential is missing.
- An IF node routes to Jira or Asana.
- Respond to Webhook returns the created task’s ID and anything the workflow had to guess.
Prerequisites
- n8n on Cloud or self-hosted, with the instance reachable over HTTPS if clients are outside your network. See deploying n8n on a Google Cloud VM if you need a server.
- An OpenAI API key.
- Jira Cloud with an API token and/or Asana with a personal access token, and the IDs described below.
- Audio in a supported format: OpenAI’s speech-to-text guide lists mp3, mp4, mpeg, mpga, m4a, wav and webm, up to 25 MB per file. Phone voice memos (m4a) work as they are. The old “normalise to 16 kHz mono WAV” advice isn’t needed.
Choosing a Transcription Model
OpenAI’s current transcription options, per its documentation, are gpt-transcribe (listed as the recommended general model), gpt-4o-transcribe, gpt-4o-mini-transcribe, gpt-4o-transcribe-diarize (labels who spoke) and whisper-1, now described as legacy (it is the one that supports timestamps and subtitles). The API accepts a prompt parameter for context about the recording, which helps with jargon such as project keys, and the docs note that the timestamp_granularities option works only with whisper-1. For dictated tasks you don’t need timestamps, so I use gpt-transcribe. Model names and prices change often, so check the page, and compare accuracy on a few of your own recordings before committing.
Why not n8n’s OpenAI node? It has a “Transcribe a Recording” action, which is the easiest route if whisper-1 is good enough for you. In n8n 2.41.4 it sends whisper-1 and has no model option, so to use a newer model the workflow below calls the transcription endpoint with an HTTP Request node instead, using your same OpenAI credential. If a later n8n version adds a model choice, switch back to the node.
Step 1: Collect the IDs You Will Need
- Jira: n8n’s Jira node selects projects and issue types by numeric ID (list or ID mode), not by the key you speak. Find a project’s ID by opening
https://YOUR-SITE.atlassian.net/rest/api/3/project/ENGwhile logged in (the JSON contains"id"), and find the issue type ID in Jira’s issue-type settings or from/rest/api/3/issuetype. - Asana: the workspace ID, which the Asana node can list when you pick it from its dropdown.
- A webhook secret: in n8n create a Header Auth credential (for example name
X-Webhook-Secret, value a long random string).
Step 2: Import the Workflow
Paste the JSON into a new n8n workflow (Import from clipboard) and set your credentials: the header-auth secret on the Webhook, your OpenAI key on the transcription and chat model nodes, and your Jira and Asana credentials. Then edit the lookup values in the validation node (next section) and the Asana workspace ID.
{
"name": "Voice note to Jira or Asana task",
"nodes": [
{
"parameters": {
"httpMethod": "POST",
"path": "voice-to-action",
"authentication": "headerAuth",
"responseMode": "responseNode",
"options": {}
},
"name": "Receive voice note",
"type": "n8n-nodes-base.webhook",
"typeVersion": 2.1,
"position": [
0,
0
],
"webhookId": "00000000-0000-0000-0000-000000000021",
"credentials": {
"httpHeaderAuth": {
"id": "REPLACE",
"name": "Voice webhook secret"
}
}
},
{
"parameters": {
"method": "POST",
"url": "https://api.openai.com/v1/audio/transcriptions",
"authentication": "predefinedCredentialType",
"nodeCredentialType": "openAiApi",
"sendBody": true,
"contentType": "multipart-form-data",
"bodyParameters": {
"parameters": [
{
"parameterType": "formBinaryData",
"name": "file",
"inputDataFieldName": "audio"
},
{
"name": "model",
"value": "gpt-transcribe"
},
{
"name": "prompt",
"value": "Dictated work tasks. Mentions Jira or Asana, project keys such as ENG, and people by first name."
}
]
},
"options": {}
},
"name": "Transcribe with OpenAI",
"type": "n8n-nodes-base.httpRequest",
"typeVersion": 4.5,
"position": [
240,
0
],
"retryOnFail": true,
"maxTries": 3,
"waitBetweenTries": 2000,
"credentials": {
"openAiApi": {
"id": "REPLACE",
"name": "OpenAI account"
}
}
},
{
"parameters": {
"text": "={{ $json.text }}",
"schemaType": "manual",
"inputSchema": "{\n \"type\": \"object\",\n \"properties\": {\n \"platform\": {\n \"type\": \"string\",\n \"enum\": [\n \"jira\",\n \"asana\"\n ],\n \"description\": \"Where to create the task. Use jira if the speaker says Jira, asana if they say Asana.\"\n },\n \"project\": {\n \"type\": \"string\",\n \"description\": \"Jira project key such as ENG, or the Asana project name, exactly as spoken. Empty if not mentioned.\"\n },\n \"summary\": {\n \"type\": \"string\",\n \"description\": \"Short imperative task title, under 100 characters\"\n },\n \"description\": {\n \"type\": \"string\",\n \"description\": \"Any extra detail the speaker gave\"\n },\n \"priority\": {\n \"type\": \"string\",\n \"enum\": [\n \"low\",\n \"medium\",\n \"high\"\n ],\n \"description\": \"Only if the speaker stated a priority\"\n },\n \"assignee_name\": {\n \"type\": \"string\",\n \"description\": \"The spoken name of the assignee, if any\"\n }\n },\n \"required\": [\n \"platform\",\n \"summary\"\n ]\n}",
"options": {
"systemPromptTemplate": "Extract a work task from this dictated voice note. Do not invent values that were not spoken."
}
},
"name": "Extract task details",
"type": "@n8n/n8n-nodes-langchain.informationExtractor",
"typeVersion": 1.2,
"position": [
480,
0
]
},
{
"parameters": {
"model": {
"__rl": true,
"mode": "id",
"value": "gpt-6-luna"
},
"options": {
"temperature": 0
}
},
"name": "OpenAI Chat Model",
"type": "@n8n/n8n-nodes-langchain.lmChatOpenAi",
"typeVersion": 1.3,
"position": [
480,
220
],
"credentials": {
"openAiApi": {
"id": "REPLACE",
"name": "OpenAI account"
}
}
},
{
"parameters": {
"jsCode": "// Checks the extractor's output, fills safe defaults and records what we had to guess.\n// n8n's Jira node wants numeric IDs, not the project key you say out loud. Map spoken keys to IDs here\n// (find an ID at https://YOUR-SITE.atlassian.net/rest/api/3/project/ENG), or look them up with an HTTP Request node.\nconst JIRA_PROJECTS = { ENG: '10000', OPS: '10001' }; // <- replace with your own projects\nconst JIRA_ISSUE_TYPE_ID = '10002'; // <- the ID of \"Task\" (or \"Bug\") in your Jira\nconst x = $json.output || {};\nconst transcript = $('Transcribe with OpenAI').item.json.text || '';\nconst problems = [];\nconst guessed = [];\n\nlet platform = String(x.platform || '').toLowerCase();\nif (!['jira', 'asana'].includes(platform)) { platform = 'jira'; guessed.push('platform (defaulted to jira)'); }\n\nconst summary = String(x.summary || '').trim();\nif (!summary) problems.push('no task summary was understood');\nconst projectKey = String(x.project || '').trim().toUpperCase();\nif (platform === 'jira' && !JIRA_PROJECTS[projectKey]) problems.push('Jira project \"' + projectKey + '\" is not one of: ' + Object.keys(JIRA_PROJECTS).join(', '));\n\nif (problems.length) {\n throw new Error('Cannot create task: ' + problems.join('; ') + '. Transcript: \"' + transcript.slice(0, 200) + '\"');\n}\n\nreturn { json: {\n platform, guessed,\n project: x.project || null,\n jiraProjectId: JIRA_PROJECTS[projectKey] || null,\n jiraIssueTypeId: JIRA_ISSUE_TYPE_ID,\n summary: summary.slice(0, 250), // Jira summaries are limited to 255 characters\n description: (x.description || '') + '\\n\\n---\\nCreated from a voice note. Transcript:\\n' + transcript,\n priority: x.priority || null,\n assignee_name: x.assignee_name || null,\n} };\n"
},
"name": "Validate and fill defaults",
"type": "n8n-nodes-base.code",
"typeVersion": 2,
"position": [
800,
0
]
},
{
"parameters": {
"conditions": {
"options": {
"caseSensitive": true,
"leftValue": "",
"typeValidation": "strict",
"version": 2
},
"conditions": [
{
"id": "is-jira",
"leftValue": "={{ $json.platform }}",
"rightValue": "jira",
"operator": {
"type": "string",
"operation": "equals"
}
}
],
"combinator": "and"
},
"options": {}
},
"name": "Jira or Asana?",
"type": "n8n-nodes-base.if",
"typeVersion": 2.3,
"position": [
1040,
0
]
},
{
"parameters": {
"resource": "issue",
"operation": "create",
"project": {
"__rl": true,
"mode": "id",
"value": "={{ $json.jiraProjectId }}"
},
"issueType": {
"__rl": true,
"mode": "id",
"value": "={{ $json.jiraIssueTypeId }}"
},
"summary": "={{ $json.summary }}",
"additionalFields": {
"description": "={{ $json.description }}"
}
},
"name": "Create Jira issue",
"type": "n8n-nodes-base.jira",
"typeVersion": 1,
"position": [
1280,
-120
],
"credentials": {
"jiraSoftwareCloudApi": {
"id": "REPLACE",
"name": "Jira Cloud"
}
}
},
{
"parameters": {
"resource": "task",
"operation": "create",
"workspace": "REPLACE_WITH_WORKSPACE_ID",
"name": "={{ $json.summary }}",
"otherProperties": {
"notes": "={{ $json.description }}"
}
},
"name": "Create Asana task",
"type": "n8n-nodes-base.asana",
"typeVersion": 1,
"position": [
1280,
120
],
"credentials": {
"asanaApi": {
"id": "REPLACE",
"name": "Asana token"
}
}
},
{
"parameters": {
"respondWith": "json",
"responseBody": "={{ { ok: true, platform: $('Validate and fill defaults').item.json.platform, id: $json.key || $json.gid, summary: $('Validate and fill defaults').item.json.summary, guessed: $('Validate and fill defaults').item.json.guessed } }}",
"options": {}
},
"name": "Respond to caller",
"type": "n8n-nodes-base.respondToWebhook",
"typeVersion": 1.5,
"position": [
1520,
0
]
}
],
"connections": {
"Receive voice note": {
"main": [
[
{
"node": "Transcribe with OpenAI",
"type": "main",
"index": 0
}
]
]
},
"Transcribe with OpenAI": {
"main": [
[
{
"node": "Extract task details",
"type": "main",
"index": 0
}
]
]
},
"Extract task details": {
"main": [
[
{
"node": "Validate and fill defaults",
"type": "main",
"index": 0
}
]
]
},
"OpenAI Chat Model": {
"ai_languageModel": [
[
{
"node": "Extract task details",
"type": "ai_languageModel",
"index": 0
}
]
]
},
"Validate and fill defaults": {
"main": [
[
{
"node": "Jira or Asana?",
"type": "main",
"index": 0
}
]
]
},
"Jira or Asana?": {
"main": [
[
{
"node": "Create Jira issue",
"type": "main",
"index": 0
}
],
[
{
"node": "Create Asana task",
"type": "main",
"index": 0
}
]
]
},
"Create Jira issue": {
"main": [
[
{
"node": "Respond to caller",
"type": "main",
"index": 0
}
]
]
},
"Create Asana task": {
"main": [
[
{
"node": "Respond to caller",
"type": "main",
"index": 0
}
]
]
}
},
"settings": {
"executionOrder": "v1"
}
}
What each node does
- Receive voice note (Webhook): POST on
/voice-to-action, authenticated with the Header Auth credential, set to answer from a Respond to Webhook node at the end, so the caller gets one definitive reply. The audio arrives as binary data under the name of the form field, hereaudio. - Transcribe with OpenAI (HTTP Request): a multipart POST to OpenAI’s
/v1/audio/transcriptionswith the file, the model, and a shortpromptthat primes the model with the vocabulary of task dictation. Retries are on (three tries, two seconds apart) for transient errors. - Extract task details: n8n’s Information Extractor fills a JSON schema. Platform and priority are limited to allowed values, and the instruction says not to invent anything that wasn’t said. The chat model is set to temperature 0 for consistency;
gpt-6-lunais the efficiency-focused model on OpenAI’s models page, and a small model is enough here. Change it to any current model. - Validate and fill defaults: the Code node below.
- Jira or Asana? routes on the platform field.
- Create Jira issue / Create Asana task: create the task with the summary and a description that includes the full transcript, so there is an audit trail of what was actually said.
- Respond to caller: returns
ok, the platform, the new key or ID, the summary and a list of things the workflow guessed.
The Validation Code
Language models sometimes return something unusable, and callers sometimes dictate something ambiguous. This node is the guard rail: it checks that a summary exists, defaults the platform to Jira if none was said (and records that it guessed), maps the spoken Jira project key to its ID, and throws a descriptive error otherwise. I ran it against sample extractor outputs: Asana with a project name (passes), Jira with a lowercase key (mapped to the right ID), an empty summary (error), an unknown project (error naming the allowed ones), and no project (error).
// Checks the extractor's output, fills safe defaults and records what we had to guess.
// n8n's Jira node wants numeric IDs, not the project key you say out loud. Map spoken keys to IDs here
// (find an ID at https://YOUR-SITE.atlassian.net/rest/api/3/project/ENG), or look them up with an HTTP Request node.
const JIRA_PROJECTS = { ENG: '10000', OPS: '10001' }; // <- replace with your own projects
const JIRA_ISSUE_TYPE_ID = '10002'; // <- the ID of "Task" (or "Bug") in your Jira
const x = $json.output || {};
const transcript = $('Transcribe with OpenAI').item.json.text || '';
const problems = [];
const guessed = [];
let platform = String(x.platform || '').toLowerCase();
if (!['jira', 'asana'].includes(platform)) { platform = 'jira'; guessed.push('platform (defaulted to jira)'); }
const summary = String(x.summary || '').trim();
if (!summary) problems.push('no task summary was understood');
const projectKey = String(x.project || '').trim().toUpperCase();
if (platform === 'jira' && !JIRA_PROJECTS[projectKey]) problems.push('Jira project "' + projectKey + '" is not one of: ' + Object.keys(JIRA_PROJECTS).join(', '));
if (problems.length) {
throw new Error('Cannot create task: ' + problems.join('; ') + '. Transcript: "' + transcript.slice(0, 200) + '"');
}
return { json: {
platform, guessed,
project: x.project || null,
jiraProjectId: JIRA_PROJECTS[projectKey] || null,
jiraIssueTypeId: JIRA_ISSUE_TYPE_ID,
summary: summary.slice(0, 250), // Jira summaries are limited to 255 characters
description: (x.description || '') + '\n\n---\nCreated from a voice note. Transcript:\n' + transcript,
priority: x.priority || null,
assignee_name: x.assignee_name || null,
} };
Two things to note. First, the error stops the run, and since the Webhook is set to respond from a node, a failed run would otherwise leave the caller waiting; set an Error Workflow (a separate workflow that starts with an Error Trigger node) that alerts you in Slack, and consider a second Respond to Webhook branch for failures. Second, I deliberately don’t pass priority or assignee to Jira: Jira expects priority names and account IDs that vary by site, so add a small lookup (a map like the project one, or a call to Jira’s user search) before wiring those fields in.
Step 3: Call It
Use the production URL (shown on the Webhook node once the workflow is active; it contains /webhook/, while the test URL contains /webhook-test/). Send the header from your credential and the audio as form data:
curl -X POST "https://n8n.example.com/webhook/voice-to-action" \
-H "X-Webhook-Secret: YOUR_LONG_RANDOM_SECRET" \
-F "[email protected]"
A successful reply looks like {"ok":true,"platform":"jira","id":"ENG-123","summary":"Fix login redirect","guessed":[]}. On a phone, an iOS Shortcut or Android automation that records audio and POSTs it to this URL gives you a one-tap “dictate a task” button.
Testing
- Start with a short clip that names the platform, project and what to do: “Create a Jira task in ENG to update the pricing page copy.”
- Open the execution in n8n and check each stage: the transcript text, then the extracted fields, then the created task.
- Try failure cases on purpose: no project spoken, an unknown project, an empty or silent file, a file over 25 MB. Each should produce a clear error, not a half-created task.
- Try noisier recordings and short names. If accuracy drops, add vocabulary to the transcription
prompt(project names, colleagues), or compare another model. OpenAI’s guidance for splitting long audio is to avoid cutting mid-sentence, since that removes context. - Review the first twenty or so created tasks by hand. Spoken input is ambiguous, and you want to see the real mistakes before people rely on it.
Security, Privacy and Cost
- Secret header. Without it, anyone who finds the URL could create tasks and spend your OpenAI credit. Keep the secret out of client code that ships to users where possible, and rotate it if it leaks.
- Voice is personal data. Recordings can contain names and sensitive information, and you are sending them to OpenAI. Tell the people using it, check OpenAI’s current data-retention terms for your account, and don’t store audio longer than you need. n8n keeps binary data for executions according to its settings; with default pruning (14 days of execution data), uploaded audio may sit on your server longer than you expect, so reduce the execution data you save or turn off saving successful executions in the workflow settings if that matters.
- Prompt injection. Anything in a transcript, including words spoken by someone else in the room, ends up in an LLM prompt. This workflow limits the damage because the model can only fill a fixed schema and can’t call tools, and the Code node validates the result. Keep it that way: if you later give the model tools, require confirmation.
- Cost. You pay per audio minute for transcription and per token for the extraction step, both small for short notes. Check OpenAI’s pricing page for your chosen model, and set a monthly spend limit.
- Reliability. For long recordings, split them into chunks under 25 MB at silences and join the transcripts; for meetings, a diarization model such as
gpt-4o-transcribe-diarizeseparates speakers, but turning a meeting into tasks is a harder problem than this workflow solves.
Optional Improvements
- Confirmation loop: post the parsed fields to Slack with approve and edit buttons (n8n’s Slack node has a send-and-wait operation) before creating anything. This is the single biggest accuracy and trust win. See the Slack guide for setting up a Slack app.
- More platforms: replace the IF node with a Switch node and add Trello, GitHub Issues or Linear branches.
- Self-hosted transcription: if privacy or volume demands it, run an open speech model on your own server and call it from the same HTTP Request node. Execute Command is off by default in n8n 2.0 and unavailable on n8n Cloud, so prefer an HTTP service over shelling out. n8n’s v2.0 breaking-changes page explains how to re-enable the node on a self-hosted instance if you really need it.
Which audio formats can I send to the OpenAI transcription API?
OpenAI lists mp3, mp4, mpeg, mpga, m4a, wav and webm, with files up to 25 MB. Common phone recordings such as m4a work directly, so a conversion step with ffmpeg is usually unnecessary.
Which model should I use for transcription: Whisper or a newer one?
OpenAI’s documentation now recommends gpt-transcribe for general transcription and describes whisper-1 as legacy, still useful when you need timestamps or subtitles. Test a few of your own recordings, especially with accents and jargon, and check current prices.
Why is the Execute Command node missing?
n8n 2.0 disables the Execute Command and Local File Trigger nodes by default for security reasons, and they are not available on n8n Cloud. On a self-hosted instance you can re-enable them by changing the NODES_EXCLUDE environment variable, but it is better to call an HTTP service instead.
How do I map a spoken name like “John” to a Jira or Asana user?
Keep a small lookup, either a map in a Code node or a call to the tool’s user-search API, and resolve the name to an account ID before creating the task. Treat unknown or ambiguous names as an error or leave the task unassigned rather than guessing.
Can a mobile app send audio directly to n8n?
Yes. Point the app at the Webhook’s production URL and send the audio as multipart form data with the secret header. For large files or chunked uploads, put a small gateway in front and pass n8n a URL, keeping in mind the 25 MB per-file limit of the transcription API.