Skip to content

Releases: harry0703/MoneyPrinterTurbo

v1.3.7

Choose a tag to compare

@harry0703 harry0703 released this 13 Sep 12:26

English | 中文

Highlights

  • Added word-by-word subtitles and a pop-up spring animation, with independent controls for display mode and animation in the WebUI.
  • Added preset background music selection with in-browser audio preview, so you can listen before generating a video.
  • Added native pauses to scripts with Azure TTS V1 (Edge TTS): use [pause: 2s] or [停顿: 2s] to insert silence while keeping narration and subtitles aligned.
  • Added Kokoro and VoxCPM TTS providers. Kokoro supports a self-hosted OpenAI-compatible service and automatically loads the server's voice list.
  • Added Claude Code as an LLM provider using an authenticated Claude subscription, with dedicated Docker deployment files; also added the OpenAI-compatible API Route provider.
  • Improved Shengsuan Cloud video generation with account-specific model discovery, quotes and explicit cost confirmation, plus clip-count recommendations based on narration duration.
  • Reduced repeated material selection across videos generated in the same task when using Pexels, Pixabay, Coverr, or local files. Material reuse is reported when the available pool is insufficient.
  • Added a YouTube “Made for kids” audience setting to Upload-Post publishing.
  • Added Azerbaijani WebUI localization and updated translations for the new settings and prompts.

Fixes and Improvements

  • Improved subtitle animation rendering, including synchronized scaling of subtitle images and transparency masks to avoid black outlines when subtitles appear.
  • Hardened Kokoro voice discovery and audio handling: retain the selected voice when the service is unavailable and validate generated audio before replacing existing files.
  • Added pause-tag validation and improved multi-segment narration/subtitle synchronization; invalid speech chunks now stop generation rather than silently producing incomplete narration.
  • Handled non-finite voice rates such as NaN and infinity more safely.
  • Fixed the WebUI launch script's port check when the project path contains spaces, and corrected imports and repository paths in the Colab notebook.
  • Added periodic log messages while FFmpeg concatenation is running, making long-running processing easier to diagnose.
  • Extended material cache cleanup to reclaim temporary files left behind by interrupted writes.
  • Updated the agent helper to support openai_image and recognize claude_code as a keyless LLM provider.
  • Synced the documented Edge TTS voice list with the built-in catalog and clarified where to find the Windows one-click package in release Assets.

Usage and Upgrade Notes

  • Subtitle defaults remain sentence display with no animation. Word-by-word timing works best with Azure TTS V1 (Edge TTS) or Whisper; providers without word-level timestamps may display phrases or whole sentences instead.
  • Script pause tags are supported only with Azure TTS V1 (Edge TTS), with individual pauses from 0.1 to 10 seconds.
  • Kokoro requires a separately running compatible TTS service. Claude Code requires its CLI and subscription authentication; selecting either provider does not install or authenticate the underlying service automatically.
  • The YouTube audience setting is an explicit user declaration, not automatic content classification. It defaults to not made for kids; review the setting before uploading.
  • Batch material allocation reduces reuse within a single task; it does not guarantee unique footage when there are too few source clips and does not change paid AI material providers' allocation behavior.

What's Changed

Thanks to all contributors and everyone who reported issues and helped test these changes!

Full Changelog: v1.3.6...v1.3.7


重点更新

  • 新增逐词字幕和弹跳出现动画,WebUI 可分别选择字幕显示模式与动画效果。
  • 新增预设背景音乐选择与浏览器内试听,生成视频前即可确认配乐效果。
  • 新增 Azure TTS V1(Edge TTS)脚本停顿支持:使用 [停顿: 2s] 或 [pause: 2s] 插入静音,并同步调整配音与字幕时间轴。
  • 新增 Kokoro 和 VoxCPM 配音服务。Kokoro 支持自托管的 OpenAI 兼容服务,并自动读取服务端音色列表。
  • 新增通过已认证 Claude 订阅使用的 Claude Code 大模型服务,提供专用 Docker 部署文件;同时新增 OpenAI 兼容的 API Route 大模型服务。
  • 优化胜算云视频生成流程,支持读取当前账号可用模型、获取报价和确认费用,并根据旁白时长推荐素材数量。
  • 减少同一任务批量生成多个视频时的素材重复,适用于 Pexels、Pixabay、Coverr 和本地文件;素材不足而发生复用时会给出提示。
  • Upload-Post 新增 YouTube“是否面向儿童”受众设置。
  • 新增阿塞拜疆语 WebUI 翻译,并补齐新设置与提示的多语言文案。

修复与改进

  • 优化字幕动画渲染,同步缩放字幕画面与透明蒙版,避免字幕出现时产生黑色轮廓。
  • 完善 Kokoro 音色读取与音频处理:服务暂时不可用时保留原有音色选择,替换文件前校验生成音频。
  • 增加停顿标签校验,改进分段配音与字幕同步;无效语音片段会中止生成,避免静默输出缺失旁白的成片。
  • 完善 NaN、无穷大等非有限语速值的处理。
  • 修复项目路径包含空格时 WebUI 启动脚本的端口检查,并修正 Colab notebook 的缺失导入和仓库路径。
  • FFmpeg 拼接期间增加定期日志,便于判断长时间处理状态和排查问题。
  • 素材缓存清理支持回收写入中断后遗留的临时文件。
  • Agent 辅助脚本新增 openai_image 支持,并正确识别无需 API Key 的 claude_code 服务。
  • 将 Edge TTS 文档音色列表与内置音色库同步,并明确 Windows 一键包在 Release Assets 中的下载位置。

使用与升级说明

  • 字幕默认仍为按句显示、无动画。逐词时间轴在 Azure TTS V1(Edge TTS)或 Whisper 下最准确;其他服务缺少逐词时间戳时,可能按短语或整句显示。
  • 脚本停顿标签仅支持 Azure TTS V1(Edge TTS),单次停顿范围为 0.1~10 秒。
  • Kokoro 需要单独运行兼容的配音服务;Claude Code 需要安装 CLI 并完成订阅认证。选择服务商不会自动完成底层服务的安装或认证。
  • YouTube 受众设置由用户主动声明,不会自动识别视频内容。默认非面向儿童,上传前请根据实际内容确认。
  • 批量素材分配仅减少同一任务内的重复;素材不足时仍可能复用,不保证每条视频的画面完全不同,也不改变付费 AI 素材服务的分配行为。

主要变更

Read more

v1.3.6

Choose a tag to compare

@harry0703 harry0703 released this 02 Sep 14:02

English | 中文

Highlights

  • Added native Volcano Engine Ark Seedance video generation for generating source clips directly from script segments.
  • Added OFox multi-model text-to-video support, allowing Seedance, Wan, and other models to share one API key.
  • Added Metaso MiniMax H3 text-to-video support with 768P/2K output, 4–15 second clips, and 9:16, 16:9, and 1:1 aspect ratios.
  • Added an OpenAI-compatible text-to-image material source for cloud relays and local gateways.
  • Added OpenRouter and APIMart as OpenAI-compatible LLM providers.
  • Added a WebUI Auto-Publish settings tab for Upload-Post, with clearer API key and account setup guidance.
  • Added CLI batch manifest mode and applied saved WebUI settings as CLI defaults when command-line values are omitted.
  • Added configurable source clip fit modes: cover fills the frame by cropping excess edges, while contain preserves the complete source with letterboxing.
  • Added complete French and Korean WebUI translations.
  • Reorganized material source settings, grouped video sources by capability, and centralized AI video provider credentials in the WebUI settings.

Security

  • Restricted browser cross-origin API access to same-origin requests by default. Trusted standalone browser frontends can be enabled explicitly with CORS_ALLOWED_ORIGINS.
  • Added provenance and SBOM attestations to published container images.

Fixes and Improvements

  • Measured narration duration from the generated audio file instead of relying only on subtitle cues.
  • Anchored the SiliconFlow subtitle timeline to the complete audio duration.
  • Ensured AudioFileClip resources are closed in ElevenLabs, Chatterbox, and Fish Audio paths.
  • Prevented script cleanup from greedily removing unrelated bracketed or parenthesized text.
  • Removed a misleading retry warning after the final LLM attempt.
  • Fixed Redis URLs when no password is configured.
  • Preserved log records when the project root and log destination are on different mounts.
  • Fixed Unicode console output after successful CLI generation on Windows.
  • Added safe, portable download filenames, including RFC-compatible quoting and protection from Windows-reserved names.
  • Improved WebUI preview behavior on headless servers by falling back to browser playback and download.
  • Decoupled Upload-Post configuration from automatic publishing so settings can be saved without enabling uploads.
  • Hardened OpenAI-compatible image handling: malformed 200 responses now skip only the affected keyword, while local storage failures stop further paid requests.

Upgrade Notes

  • Browser API access now defaults to same-origin requests. Standalone web frontends must declare their trusted origins through CORS_ALLOWED_ORIGINS; server-side clients such as curl, Postman, and n8n are unaffected.
  • CLI commands now inherit saved WebUI settings for omitted options. Explicit CLI arguments and batch manifest values still take precedence.

What's Changed

Full Changelog: v1.3.5...v1.3.6


重点更新

  • 新增火山引擎方舟 Seedance 原生视频生成服务,可根据脚本片段直接生成视频素材。
  • 新增 OFox 多模型文生视频服务,一个 API Key 即可调用 Seedance、Wan 等模型。
  • 新增秘塔 MiniMax H3 文生视频服务,支持 768P/2K、4~15 秒片段,以及 9:16、16:9 和 1:1 三种画幅。
  • 新增适用于云端中转和本地网关的 OpenAI 兼容文生图素材源。
  • 新增 OpenRouter 和 APIMart OpenAI 兼容大语言模型服务商。
  • WebUI 新增 Upload-Post 自动发布设置,并完善 API Key 和账号名称的申请、配置指引。
  • CLI 新增批量清单模式;未显式传入命令行参数时,可以继承 WebUI 已保存的默认设置。
  • 新增素材画面适配模式:cover 通过裁剪边缘填满画面,contain 保留完整素材并在比例不同时增加黑边。
  • 新增完整的法语和韩语 WebUI 翻译。
  • 重新整理素材来源设置,按能力对视频来源分组,并将 AI 视频服务商凭据统一集中到 WebUI 设置中管理。

安全性

  • 浏览器跨域 API 访问现在默认遵循同源策略;如需使用独立网页前端,可以通过 CORS_ALLOWED_ORIGINS 显式配置可信来源。
  • 发布的容器镜像新增构建来源证明和 SBOM 软件物料清单证明。

修复与改进

  • 改为从真实生成的音频文件测量旁白时长,不再仅依赖字幕时间轴。
  • 将 SiliconFlow 字幕时间轴末尾与完整音频时长对齐。
  • 确保 ElevenLabs、Chatterbox 和 Fish Audio 链路始终关闭 AudioFileClip 资源。
  • 修复脚本文本清理可能贪婪删除多组括号内容的问题。
  • 移除大语言模型最后一次重试结束后误导性的“正在重试”日志。
  • 修复 Redis 未配置密码时生成错误鉴权片段的问题。
  • 修复项目目录和日志目录位于不同挂载点时日志记录丢失的问题。
  • 修复 Windows 终端无法编码 Unicode 时,CLI 在视频成功生成后异常退出的问题。
  • 改进下载文件名的安全性和跨平台兼容性,包括标准引用编码以及 Windows 保留名称保护。
  • 无图形桌面的服务器环境下,WebUI 会回退到浏览器播放和下载视频。
  • 将 Upload-Post 配置与自动发布开关解耦,允许用户保存配置但不立即启用上传。
  • 加强 OpenAI 兼容文生图异常处理:HTTP 200 返回损坏图片时仅跳过当前关键词,本地存储失败时则停止后续付费请求。

升级说明

  • 浏览器调用 API 现在默认遵循同源策略。独立网页前端需要通过 CORS_ALLOWED_ORIGINS 配置可信来源;curl、Postman、n8n 等服务端调用不受影响。
  • CLI 未显式传入参数时,会继承 WebUI 保存的设置;明确传入的 CLI 参数和批量清单字段仍拥有更高优先级。

主要变更

  • #1246:在英文和日语文档中补充 Fish Audio,感谢 @c020627。
  • #1140:为容器镜像增加构建来源证明和 SBOM,感谢 @kobihikri。
  • #1209:新增 Upload-Post 自动发布 WebUI 设置,感谢 @parveen0029。
  • #1252:改进无图形桌面服务器上的 WebUI 视频预览,感谢 @abhiunix。
  • #1225:CLI 支持继承 WebUI 已保存设置,感谢 @Gainto。
  • #1259:将 Upload-Post 配置和自动发布解耦,感谢 @mcgeand。
  • #1260:修复跨文件系统挂载时日志记录丢失,感谢 @SandroHub013。
  • #1243:新增 CLI 批量清单模式,感谢 @lihuiyang1024。
  • #1265:修复 API 下载文件名引用编码,感谢 @Mihir7027。
  • #1264:使用更易读的批量清单大小错误提示,感谢 @Mihir7027。
  • #1266:修复 Redis 未配置密码时的连接 URL,感谢 @Mihir7027。
  • #1267:移除最后一次 LLM 尝试后的错误重试提示,感谢 @Mihir7027。
  • #1268:将 SiliconFlow 字幕与完整音频时长对齐,感谢 @Mihir7027。
  • #1263:从真实音频文件测量旁白时长,感谢 @YUSAKRU。
  • #1269:可靠释放 TTS 音频资源,感谢 @Mihir7027。
  • #1270:修复脚本括号内容清理的贪婪匹配,感谢 @Mihir7027。
  • ab1c790:新增 APIMart OpenAI 兼容大语言模型服务商。
  • #1277:新增 OpenRouter 大语言模型服务商,感谢 @brizzio。
  • #1282:修复部分 Windows 终端...
Read more

v1.3.5

Choose a tag to compare

@harry0703 harry0703 released this 22 Aug 14:29

English | 中文

Highlights

  • Added MiniMax and Fish Audio text-to-speech providers, improved MiniMax regional endpoint and voice selection support, synchronized Gemini voices with the official catalog, and added automatic voice previews in the WebUI.
  • Added Anthropic Claude as a native LLM provider.
  • Added Shengsuan AI and WaveSpeed text-to-video material sources.
  • Added reusable WebUI generation settings and preset import/export, including optional API key backup and restore.
  • Added LLM-generated social captions for every supported cross-post platform, instead of limiting them to YouTube.
  • Added an Italian WebUI translation and a Japanese README.

Security

  • Added optional API key authentication for /api/v1 endpoints and generated task files. Existing installations remain backward compatible while app.api_key is empty.
  • Restricted custom audio inputs to allowed locations to prevent path traversal.
  • Prevented generated task files from being accessed through escaping symlinks.
  • Hardened local material uploads with stricter file and path validation.
  • Sanitized client-supplied request IDs before using them in task processing.

Fixes and Improvements

  • Fixed the last subtitle line being clipped for some text and font combinations, including multiline subtitles and text without descenders.
  • Added an actionable startup error when FFmpeg is unavailable.
  • Fixed Docker dependency installation retries so persistent package failures stop the build instead of being silently ignored.
  • Rolled back task state when background scheduling fails, preventing tasks from remaining stuck in an incorrect state.
  • Improved API response schemas and migrated model configuration to Pydantic v2 ConfigDict.
  • Registered the ping health-check router.
  • Preserved paragraph boundaries in normalized LLM responses.
  • Distinguished the supported Kimi API platforms in configuration and the WebUI.
  • Handled configuration files containing repeated UTF-8 BOM markers.
  • Improved MiniMax voice discovery and validated generated audio before replacing an existing output file.

Upgrade Notes

  • No authentication is required by default. To protect the V1 API and generated task files, set app.api_key in config.toml and send the same value in the x-api-key request header.
  • Settings preset exports can contain provider credentials when API key backup is enabled. Store exported preset files securely and do not commit them to source control.
  • Fish Audio, MiniMax, Anthropic Claude, WaveSpeed, and Shengsuan AI require their respective provider credentials before they can be used.

What's Changed

New Contributors

Full Changelog: v1.3.4...v1.3.5


重点更新

  • 新增 MiniMax 和 Fish Audio 文本转语音服务,完善 MiniMax 区域端点和音色选择支持,同步 Gemini 官方音色目录,并在 WebUI 中增加自动语音试听功能。
  • 新增 Anthropic Claude 原生大语言模型服务商。
  • 新增胜算云(Shengsuan AI)和 WaveSpeed AI 文生视频素材源。
  • WebUI 新增可复用的视频生成设置以及预设导入、导出功能,并支持选择性备份和恢复 API Key。
  • 所有受支持的跨平台发布渠道现在都可以使用大语言模型生成社交媒体文案,不再仅限于 YouTube。
  • 新增意大利语 WebUI 翻译和日语 README 文档。

安全性

  • 为 /api/v1 接口和生成的任务文件增加可选的 API Key 鉴权;当 app.api_key 保持为空时,现有安装和客户端仍可继续使用,无需修改。
  • 限制自定义音频文件只能从允许的目录读取,防止路径穿越。
  • 防止通过符号链接访问任务目录之外的文件。
  • 加强本地素材上传的文件类型、路径和文件名校验。
  • 对客户端传入的请求 ID 进行清理和规范化处理,避免其影响任务文件路径和处理流程。

修复与改进

  • 修复部分文字和字体组合下字幕最后一行被裁切的问题,覆盖多行字幕以及不包含下伸字符的文本。
  • 当系统缺少 FFmpeg 时立即停止启动,并提供清晰、可执行的错误提示。
  • 修复 Docker 依赖安装重试逻辑,确保软件包持续安装失败时终止构建,而不是静默忽略错误。
  • 后台任务调度失败时自动回滚任务状态,避免任务停留在错误状态。
  • 改进 API 响应模型,并将相关 Pydantic 配置迁移到 Pydantic v2 ConfigDict。
  • 正确注册 Ping 健康检查接口。
  • 保留大语言模型响应中的段落和换行结构。
  • 在配置和 WebUI 中明确区分不同的 Kimi API 平台。
  • 支持读取包含多个 UTF-8 BOM 标记的配置文件。
  • 改进 MiniMax 音色发现与选择,并在替换已有音频文件前验证新生成的音频是否有效。

升级说明

  • 默认情况下不会启用 API 鉴权,现有用户升级后不受影响。如需保护 V1 API 和生成的任务文件,请在 config.toml 中设置 app.api_key,并在请求中通过 x-api-key 请求头传递相同的值。
  • 开启 API Key 备份后,导出的 WebUI 设置预设可能包含服务商凭据。请妥善保存导出文件,不要将其提交到公开代码仓库。
  • Fish Audio、MiniMax、Anthropic Claude、WaveSpeed 和胜算云功能需要配置对应服务商的凭据后才能使用。

主要变更

新贡献者

感谢 @HaningZS、@AZEROSTART、@jxnding、@eltociear、@NafeesMadni、@SandroHub013、@rignaneseleo、@chengzeyi、@Gainto、@alvinhui、@lihuiyang1024 和 @hariom-hp 首次参与 MoneyPrinterTurbo 项目贡献。

完整变更记录:v1.3.4...v1.3.5

v1.3.4

Choose a tag to compare

@harry0703 harry0703 released this 12 Aug 02:51

Highlights

  • Added a WebUI preference to disable automatically opening the task folder after video generation.
  • Generated videos now use the video subject as the download filename, with automatic numbering for multi-video tasks.
  • Added configurable Whisper initial_prompt support to improve subtitle recognition for names, terminology, and language-specific vocabulary.
  • Added persistent material search caching to reduce duplicate provider requests, API usage, and rate-limit pressure.
  • Stock material selection now prioritizes assets matching the target video orientation.
  • Added sanitized material source provenance to task artifacts for easier source tracking.

Fixes and Improvements

  • Fixed WebUI configuration changes potentially blocking the page while a video generation task was running.
  • Fixed the ElevenLabs API Key being cleared after a WebUI restart or browser reconnection.
  • Unified ElevenLabs API Key resolution across TTS and music generation, including ELEVENLABS_API_KEY environment variable support.
  • Fixed configuration updates when config.toml is provided through a bind mount.
  • Improved diagnostics for Pixabay 429 responses and Cloudflare challenges.
  • Prevented Pixabay API Keys from appearing in application logs.
  • Added response and resource limits to ElevenLabs music requests to prevent excessive memory usage.
  • API requests now reject video counts and clip durations below 1.
  • Fixed Worker crashes caused by stale or invalid queued task parameters from earlier versions.
  • Expanded the WebUI video subject input for easier editing of longer subjects.
  • Updated project documentation and related resources.

What's Changed

New Contributors

Full Changelog: v1.3.3...v1.3.4

v1.3.3

Choose a tag to compare

@harry0703 harry0703 released this 24 Jul 04:38

Highlights

  • Added full voiceover preview with estimated duration before video generation. Preview audio can be reused when the final generation settings remain unchanged.
  • Added optional video-matched background music generation through Sonilo and ElevenLabs, with graceful fallback when music generation fails.
  • Added custom background music uploads directly from the WebUI.
  • Added a global video clip speed control and new ZoomIn / ZoomOut transitions.
  • Added lightweight, non-blocking update notifications in the WebUI.
  • Hardened task execution and recovery across the WebUI, API, and CLI, with clearer stage and error tracking.

Fixes and Improvements

  • Fixed Windows portable-package startup failures caused by conflicting Python app packages.
  • Fixed the WebUI becoming blank after starting video generation.
  • Preserved generation progress, logs, and terminal task states across Streamlit reruns.
  • Improved interrupted-task recovery and kept social publishing failures independent from successful video generation.
  • Added queue limits and more reliable state updates for cross-platform publishing.
  • Prevented unexpected Whisper model fallback when another subtitle provider fails.
  • Fixed LLM provider errors being passed into material keyword generation.
  • Fixed repeated cleanup attempts for the same temporary video file.
  • Allowed video materials that fall only a few pixels below the nominal minimum resolution because of encoder rounding.
  • Standardized TTS API key links and provider setup guidance.
  • Completed missing secondary translations across supported WebUI languages.
  • Added line-ending normalization for more consistent contributions across Windows, macOS, and Linux.
  • Expanded CI coverage for task execution, API controllers, media handling, queues, application startup, and failure recovery.

What's Changed

New Contributors

Full Changelog: v1.3.2...v1.3.3

v1.3.2

Choose a tag to compare

@harry0703 harry0703 released this 12 Jul 08:54

Highlights

  • Redesigned the WebUI with a cleaner responsive layout, a dedicated settings dialog, improved dark mode support, and a more consistent interaction experience.
  • Added task management with status filtering, progress tracking, video playback, folder access, deletion, and regeneration from historical task settings.
  • Added a lightweight onboarding tour to guide new users through configuration, video options, and generation.
  • Added built-in LLM connection testing so provider credentials and model settings can be verified before generation.
  • Centralized LLM provider metadata, defaults, configuration tips, and registration links for more consistent setup and maintenance.
  • Added video cache statistics and cleanup controls.
  • Expanded the command-line workflow and added the MoneyPrinterTurbo Agent skill for automated video generation.
  • Updated the project version to v1.3.2.

Fixes and Improvements

  • Migrated Gemini text and TTS integrations from the deprecated google-generativeai SDK to the new google-genai SDK.
  • Fixed WebUI selectors that sometimes required two selections before a change took effect.
  • Preserved generation logs across Streamlit reruns and fixed logger cleanup errors after successful generation.
  • Improved task state normalization so filtering works consistently across all UI languages.
  • Prevented running tasks from being deleted and tightened task file operation boundaries.
  • Improved audio settings with clearer narration modes, provider-specific setup guidance, and more consistent preview controls.
  • Fixed Azure TTS V2 preview generation so the selected speech rate is applied correctly.
  • Improved subtitle settings, background controls, contrast warnings, default restoration, and multilingual rendering reliability.
  • Disabled subtitle backgrounds by default while retaining optional rounded translucent backgrounds.
  • Added automatic video encoder selection with safer fallback to libx264.
  • Improved video duration handling with a small safety margin to prevent incomplete output.
  • Fixed duration detection for custom audio formats other than MP3.
  • Improved LLM provider defaults, descriptions, registration links, and multilingual configuration guidance.
  • Removed deprecated or unsupported provider and video source options.
  • Improved WebUI performance by caching font and background music discovery.
  • Updated Streamlit and related dependencies, Docker configurations, startup scripts, and toolbar behavior.
  • Refined multilingual UI copy and layout across all supported languages.
  • Updated README documentation, screenshots, gallery previews, configuration examples, and the Colab notebook.
  • Expanded regression coverage for WebUI state, task management, providers, CLI workflows, subtitles, caching, and configuration.

What's Changed

Full Changelog: v1.3.1...v1.3.2

v1.3.1

Choose a tag to compare

@harry0703 harry0703 released this 06 Jul 02:34

Highlights

  • Added new LLM providers and integrations, including AIML API, EvoLink, and VolcEngine Ark.
  • Added new TTS options, including ElevenLabs and self-hosted OpenAI-compatible Chatterbox TTS.
  • Added optional TwelveLabs video AI integration for smarter material understanding and reranking.
  • Added YouTube Shorts publishing support through Upload-Post.
  • Added a pure CLI workflow and exposed more WebUI-backed video parameters to command-line generation.
  • Added GHCR Docker image publishing for easier container deployment.
  • Improved multilingual WebUI support, including Indonesian and Spanish updates, plus explicit es-ES script generation support.

Fixes and Improvements

  • Fixed subtitle background clipping and improved subtitle rendering reliability.
  • Fixed custom-audio subtitle generation when using Whisper.
  • Fixed Windows MoviePy temporary-audio handling to avoid Windows Defender interference.
  • Fixed Streamlit first-run prompt blocking WebUI startup.
  • Fixed LLM provider selectbox switching that sometimes required selecting twice.
  • Fixed video duration safety margin handling.
  • Fixed video_contact_mode typo compatibility by correctly using video_concat_mode.
  • Fixed Pollinations default API endpoint.
  • Improved robust JSON parsing for fenced or multi-line LLM responses.
  • Exposed match_materials_to_script in the terms API.
  • Improved material matching to better follow script order.
  • Hardened CLI argument validation for local video material workflows.
  • Redacted LLM error output to avoid leaking sensitive details.
  • Updated sponsor acknowledgements and provider copy.

What's Changed

New Contributors

Full Changelog: v1.3.0...v1.3.1

v1.3.0

Choose a tag to compare

@harry0703 harry0703 released this 10 Jun 09:39

Highlights

  • Added Coverr as a new stock video material provider, including WebUI API key management and 16:9-friendly defaults.
  • Added Groq as a new LLM provider.
  • Added no-voice video generation mode for workflows that do not need narration.
  • Added configurable video encoder selection and safer hardware-encoder fallback behavior.
  • Added AIHubMix provider integration.
  • Added multilingual social metadata generation API.
  • Added subtitle background options, including optional rounded translucent subtitle backgrounds.
  • Improved WebUI localization coverage, including Arabic, German, Russian, Turkish, and Vietnamese updates.

Fixes and Improvements

  • Improved repeated video material reuse behavior.
  • Fixed final SRT subtitle block handling.
  • Fixed subtitle matching for markdown-style script markers.
  • Fixed Qwen response parsing edge cases.
  • Upgraded and hardened selected dependencies and deployment defaults.
  • Refactored Azure voice data into JSON for easier maintenance.
  • Improved README and setup documentation quality.

What's Changed

New Contributors

Validation

  • uv lock --check
  • uv run --with pytest pytest test/services -q
  • Result: 117 passed, 1 skipped, 18 warnings, 9 subtests passed.

Full Changelog: v1.2.9...v1.3.0

v1.2.9

Choose a tag to compare

@harry0703 harry0703 released this 30 May 09:38

What's Changed

New features

  • Added advanced script settings in WebUI:
    • script paragraph count
    • custom script requirements
    • optional full custom system prompt for advanced users
  • Added Xiaomi MiMo LLM provider support.
  • Added Xiaomi MiMo TTS provider support.
  • Improved Ollama default base URL handling for Docker environments.

Fixes and improvements

  • Fixed custom BGM path handling for WebUI and API flows.
  • Fixed Windows WebUI launcher robustness.
  • Fixed webui.sh to bind/open 127.0.0.1 by default instead of opening 0.0.0.0.
  • Fixed string video concat mode handling in downloads.
  • Fixed subtitle wrapping for long text.
  • Fixed Edge TTS signed-zero rate formatting.
  • Added missing Turkish translation keys.
  • Fixed README-en ImageMagick code fence.

Testing

  • Added regression coverage for LLM provider behavior, local TTS/video task paths, video concat mode handling, and advanced script prompt settings.
  • Verified advanced script settings end-to-end with real script generation, TTS, Pexels material download, subtitle generation, and final video synthesis.

v1.2.8

Choose a tag to compare

@harry0703 harry0703 released this 28 May 03:44

MoneyPrinterTurbo v1.2.8

This release collects the bug fixes, provider additions, security hardening, deployment fixes, and PRs merged since v1.2.7.

Highlights

  • Added LiteLLM provider support for 100+ compatible model gateways.
  • Added Grok/xAI provider support through the existing OpenAI-compatible LLM path.
  • Added WebUI support for uploading custom audio and generating a video from local narration.
  • Improved Gemini TTS and Edge subtitle compatibility after recent dependency updates.
  • Fixed Azure LLM provider routing so AzureOpenAI requests use the Azure client path correctly.
  • Updated the Google Colab notebook to use an isolated uv environment and avoid Colab global dependency conflicts.

Bug Fixes

  • Fixed subtitle splitting so numbers with thousands separators such as 1,000 are preserved.
  • Fixed Redis task pagination and task state listing behavior.
  • Fixed bundled ffmpeg discovery for video concatenation.
  • Closed audio clips after duration probing to avoid file handle leaks.
  • Suppressed noisy MoviePy probing output during material inspection.
  • Added timeout handling for hanging Edge TTS streams.
  • Hardened LiteLLM response parsing for empty choices/messages.
  • Improved Windows portable updater and Azure TTS compatibility.
  • Restored Gemini TTS subtitle generation when using the edge subtitle provider.

Security And Hardening

  • Hardened task folder opening and path validation.
  • Hardened uploaded/retrieved media file path handling.
  • Added task queue bounds and safer queue behavior.
  • Restored TLS verification for external material/API requests.
  • Disabled risky g4f usage by default and moved it behind an explicit optional dependency path.
  • Added focused regression tests for file path, task state, LLM, material, video, and voice behavior.

Documentation And Deployment

  • Added a system requirements matrix to the README.
  • Fixed README typos in Chinese and English docs.
  • Allowed Redis host override in Docker deployments.
  • Updated Colab setup to use uv sync --frozen --python 3.11 and launch Streamlit through uv run.

Merged PRs

  • #861 Docker Redis host override.
  • #891 README typo fix.
  • #897 regression test for video_transition_mode=None.
  • #900 README-en typo fix.
  • #903 Grok provider support.

Validation

Before tagging this release:

uv lock --check
uv run python -m unittest test.services.test_llm test.services.test_video.TestVideoService.test_combine_videos_handles_none_transition_mode test.services.test_voice.TestVoiceService.test_edge_cue_aggregation_handles_thousand_separator_comma
uv run python -m compileall app webui