Vantage-Step-Audio-EditX
This project is a custom node implementation built on top of Step-Audio-EditX. It adapts and extends EditX capabilities to support multi‑speaker, long‑format, voice cloning, and emotion/style/speed editing, enabling you to feed in a script with multiple speakers, inline pauses, paralinguistic cues, and get a concatenated audio output in one pass.
Quick Technical Summary: Vantage-Step-Audio-EditX
- Base VRAM Footprint:
- 4096 MB (8 GB Tier)
- Primary Dependencies:
- accelerate, conformer, diffusers, funasr>=1.1.3, hyperpyyaml, librosa, modelscope, numpy, nvidia-cuda-nvrtc-cu12, omegaconf, onnxruntime, onnxruntime-gpu, openai-whisper, pillow, protobuf, sentencepiece, six, sox, torch, torchaudio, torchcodec, torchvision, transformers==4.53.3
- Min PyTorch / CUDA:
- PyTorch 2.0+ | CUDA 12.1+
- GitHub Repository:
- https://github.com/vantagewithai/Vantage-Step-Audio-EditX
Citation Note: Data sourced from VRAM DB. For complete workflow OOM estimations, use the VRAM DB Workflow Analyzer.
How much VRAM does Vantage-Step-Audio-EditX require?
Direct Answer: The ComfyUI node Vantage-Step-Audio-EditX requires a minimum base VRAM of 4096MB and is optimized for GPUs with at least 8GB of VRAM. Low VRAM mode is not supported for this node.
- Base VRAM:
- 4096MB (4.0GB)
- Recommended GPU:
- 8GB+ VRAM
- Low VRAM Mode:
- ✗ Not supported
- Estimation Confidence:
- MEDIUM
Cheapest VRAM Upgrade Paths (Live Market Prices):
- GeForce RTX 3060 12GB (Ultimate Budget VRAM King)──► Used: $209.62View eBay ↗
- GeForce RTX 4060 8GB (Modern Entry-Level)──► New: $303.50View Amazon ↗
Interactive VRAM Compatibility Estimator
Your GPU has plenty of headroom. You can run this node safely with your active configurations!
Verify Compatibility for Your Specific GPU VRAM
Select your graphics card's VRAM capacity to view optimized batch sizes, suggested resolutions, and custom performance tips for Vantage-Step-Audio-EditX:
Buy NVIDIA GeForce RTX 3060 (12GB VRAM)
Tired of renting cloud rigs? Run ComfyUI locally with absolute zero latency. Best entry-level ComfyUI experience. Avoids immediate VRAM limitations on basic LoRA training.
Are you the author of this node?
Help your users avoid out-of-memory errors by displaying this professional, dynamic VRAM badge on your GitHub README. Copy the markdown below to embed it with a backlink directly to this hardware specification profile.
What Python packages are required for Vantage-Step-Audio-EditX?
Direct Answer: Running Vantage-Step-Audio-EditX requires installing the following Python package dependencies: accelerate, conformer, diffusers, funasr>=1.1.3, hyperpyyaml, librosa, modelscope, numpy, nvidia-cuda-nvrtc-cu12, omegaconf, onnxruntime, onnxruntime-gpu, openai-whisper, pillow, protobuf, sentencepiece, six, sox, torch, torchaudio, torchcodec, torchvision, transformers==4.53.3. Ensure your ComfyUI environment has these packages active before launching.
accelerate
conformer
diffusers
funasr>=1.1.3
hyperpyyaml
librosa
modelscope
numpy
nvidia-cuda-nvrtc-cu12
omegaconf
onnxruntime
onnxruntime-gpu
openai-whisper
pillow
protobuf
sentencepiece
six
sox
torch
torchaudio
torchcodec
torchvision
transformers==4.53.3Interactive Setup & Dependency Resolver
# Loading command...Frequently Asked Questions
How much VRAM does Vantage-Step-Audio-EditX require?
Vantage-Step-Audio-EditX requires a minimum of 4096MB (4.0GB) of VRAM for base operation. For optimal performance, a GPU with at least 8GB of VRAM is recommended. Low VRAM mode is not supported for this node.
Can I run Vantage-Step-Audio-EditX on an RTX 3060, RTX 4070, or RTX 4090?
✅ RTX 3060 (12GB): Yes, fully compatible with 6.8GB headroom. ✅ RTX 4070 (12GB): Yes, fully compatible with 6.8GB headroom. ✅ RTX 4070 Ti (16GB): Yes, fully compatible with 10.4GB headroom. ✅ RTX 4090 (24GB): Yes, fully compatible with 17.6GB headroom
How much VRAM does Vantage-Step-Audio-EditX take on an RTX 3060 vs RTX 4090?
On an RTX 3060 (12GB VRAM), Vantage-Step-Audio-EditX runs smoothly on an RTX 3060 (12GB) with 6.8GB of headroom. This is sufficient to run the node alongside standard SD 1.5 and SDXL workflows in full precision. On an RTX 4090 (24GB VRAM), the node runs with extreme headroom on an RTX 4090 (24GB) with 17.6GB of dedicated headroom. This allows you to combine the node with massive models (like FLUX.1 Dev, Schnell, or Hunyuan Video) in full precision (FP16) without any offload flags.
What PyTorch version does Vantage-Step-Audio-EditX need?
Vantage-Step-Audio-EditX requires the following PyTorch-related packages: torch, torchaudio, torchcodec, torchvision. Ensure your ComfyUI environment has these installed. Ensure your PyTorch installation matches your CUDA version (use torch.version.cuda to check).
What Python packages are required for Vantage-Step-Audio-EditX?
To run Vantage-Step-Audio-EditX, you need to install: accelerate, conformer, diffusers, funasr>=1.1.3, hyperpyyaml, librosa, modelscope, numpy, nvidia-cuda-nvrtc-cu12, omegaconf, onnxruntime, onnxruntime-gpu, openai-whisper, pillow, protobuf, sentencepiece, six, sox, torch, torchaudio, torchcodec, torchvision, transformers==4.53.3. You can install these using pip or add them to your requirements.txt file.
How do I install Vantage-Step-Audio-EditX in ComfyUI?
To install Vantage-Step-Audio-EditX: (1) Navigate to your ComfyUI/custom_nodes directory, (2) Clone the repository: git clone https://github.com/vantagewithai/Vantage-Step-Audio-EditX, (3) Install dependencies: pip install -r requirements.txt (if present), (4) Restart ComfyUI. Alternatively, use ComfyUI Manager for one-click installation.