reflector

mirror of https://github.com/Monadical-SAS/reflector.git synced 2026-02-05 02:16:46 +00:00

Author	SHA1	Message	Date
Mathieu Virbel	9265d201b5	fix: restore previous behavior on live pipeline + audio downscaler (#561 ) This commit restore the original behavior with frame cutting. While silero is used on our gpu for files, look like it's not working great on the live pipeline. To be investigated, but at the moment, what we keep is: - refactored to extract the downscale for further processing in the pipeline - remove any downscale implementation from audio_chunker and audio_merge - removed batching from audio_merge too for now	2025-08-22 10:49:26 -06:00
Mathieu Virbel	3ea7f6b7b6	feat: pipeline improvement with file processing, parakeet, silero-vad (#540 ) * feat: improve pipeline threading, and transcriber (parakeet and silero vad) * refactor: remove whisperx, implement parakeet * refactor: make audio_chunker more smart and wait for speech, instead of fixed frame * refactor: make audio merge to always downscale the audio to 16k for transcription * refactor: make the audio transcript modal accepting batches * refactor: improve type safety and remove prometheus metrics - Add DiarizationSegment TypedDict for proper diarization typing - Replace List/Optional with modern Python list/\| None syntax - Remove all Prometheus metrics from TranscriptDiarizationAssemblerProcessor - Add comprehensive file processing pipeline with parallel execution - Update processor imports and type annotations throughout - Implement optimized file pipeline as default in process.py tool * refactor: convert FileDiarizationProcessor I/O types to BaseModel Update FileDiarizationInput and FileDiarizationOutput to inherit from BaseModel instead of plain classes, following the standard pattern used by other processors in the codebase. * test: add tests for file transcript and diarization with pytest-recording * build: add pytest-recording * feat: add local pyannote for testing * fix: replace PyAV AudioResampler with torchaudio for reliable audio processing - Replace problematic PyAV AudioResampler that was causing ValueError: [Errno 22] Invalid argument - Use torchaudio.functional.resample for robust sample rate conversion - Optimize processing: skip conversion for already 16kHz mono audio - Add direct WAV writing with Python wave module for better performance - Consolidate duplicate downsample checks for cleaner code - Maintain list[av.AudioFrame] input interface - Required for Silero VAD which needs 16kHz mono audio * fix: replace PyAV AudioResampler with torchaudio solution - Resolves ValueError: [Errno 22] Invalid argument in AudioMergeProcessor - Replaces problematic PyAV AudioResampler with torchaudio.functional.resample - Optimizes processing to skip unnecessary conversions when audio is already 16kHz mono - Uses direct WAV writing with Python's wave module for better performance - Fixes test_basic_process to disable diarization (pyannote dependency not installed) - Updates test expectations to match actual processor behavior - Removes unused pydub dependency from pyproject.toml - Adds comprehensive TEST_ANALYSIS.md documenting test suite status * feat: add parameterized test for both diarization modes - Adds @pytest.mark.parametrize to test_basic_process with enable_diarization=[False, True] - Test with diarization=False always passes (tests core AudioMergeProcessor functionality) - Test with diarization=True gracefully skips when pyannote.audio is not installed - Provides comprehensive test coverage for both pipeline configurations * fix: resolve pipeline property naming conflict in AudioDiarizationPyannoteProcessor - Renames 'pipeline' property to 'diarization_pipeline' to avoid conflict with base Processor.pipeline attribute - Fixes AttributeError: 'property 'pipeline' object has no setter' when set_pipeline() is called - Updates property usage in _diarize method to use new name - Now correctly supports pipeline initialization for diarization processing * fix: add local for pyannote * test: add diarization test * fix: resample on audio merge now working * fix: correctly restore timestamp * fix: display exception in a threaded processor if that happen * Update pyproject.toml * ci: remove option * ci: update astral-sh/setup-uv * test: add monadical url for pytest-recording * refactor: remove previous version * build: move faster whisper to local dep * test: fix missing import * refactor: improve main_file_pipeline organization and error handling - Move all imports to the top of the file - Create unified EmptyPipeline class to replace duplicate mock pipeline code - Remove timeout and fallback logic - let processors handle their own retries - Fix error handling to raise any exception from parallel tasks - Add proper type hints and validation for captured results * fix: wrong function * fix: remove task_done * feat: add configurable file processing timeouts for modal processors - Add TRANSCRIPT_FILE_TIMEOUT setting (default: 600s) for file transcription - Add DIARIZATION_FILE_TIMEOUT setting (default: 600s) for file diarization - Replace hardcoded timeout=600 with configurable settings in modal processors - Allows customization of timeout values via environment variables * fix: use logger * fix: worker process meetings now use file pipeline * fix: topic not gathered * refactor: remove prepare(), pipeline now work * refactor: implement many review from Igor * test: add test for test_pipeline_main_file * refactor: remove doc * doc: add doc * ci: update build to use native arm64 builder * fix: merge fixes * refactor: changes from Igor review + add test (not by default) to test gpu modal part * ci: update to our own runner linux-amd64 * ci: try using suggested mode=min * fix: update diarizer for latest modal, and use volume * fix: modal file extension detection * fix: put the diarizer as A100	2025-08-20 20:07:19 -06:00
Mathieu Virbel	dc177af3ff	feat: implement service-specific Modal API keys with auto processor pattern (#528 ) * fix: refactor modal API key configuration for better separation of concerns - Split generic MODAL_API_KEY into service-specific keys: - TRANSCRIPT_API_KEY for transcription service - DIARIZATION_API_KEY for diarization service - TRANSLATE_API_KEY for translation service - Remove deprecated _MODAL_API_KEY settings - Add proper validation to ensure URLs are set when using modal processors - Update README with new configuration format BREAKING CHANGE: Configuration keys have changed. Update your .env file: - TRANSCRIPT_MODAL_API_KEY → TRANSCRIPT_API_KEY - LLM_MODAL_API_KEY → (removed, use TRANSCRIPT_API_KEY) - Add DIARIZATION_API_KEY and TRANSLATE_API_KEY if using those services fix: update Modal backend configuration to use service-specific API keys - Changed from generic MODAL_API_KEY to service-specific keys: - TRANSCRIPT_MODAL_API_KEY for transcription - DIARIZATION_MODAL_API_KEY for diarization - TRANSLATION_MODAL_API_KEY for translation - Updated audio_transcript_modal.py and audio_diarization_modal.py to use modal_api_key parameter - Updated documentation in README.md, CLAUDE.md, and env.example * feat: implement auto/modal pattern for translation processor - Created TranscriptTranslatorAutoProcessor following the same pattern as transcript/diarization - Created TranscriptTranslatorModalProcessor with TRANSLATION_MODAL_API_KEY support - Added TRANSLATION_BACKEND setting (defaults to "modal") - Updated all imports to use TranscriptTranslatorAutoProcessor instead of TranscriptTranslatorProcessor - Updated env.example with TRANSLATION_BACKEND and TRANSLATION_MODAL_API_KEY - Updated test to expect TranscriptTranslatorModalProcessor name - All tests passing * refactor: simplify transcript_translator base class to match other processors - Moved all implementation from base class to modal processor - Base class now only defines abstract _translate method - Follows the same minimal pattern as audio_diarization and audio_transcript base classes - Updated test mock to use _translate instead of get_translation - All tests passing * chore: clean up settings and improve type annotations - Remove deprecated generic API key variables from settings - Add comments to group Modal-specific settings - Improve type annotations for modal_api_key parameters * fix: typing * fix: passing key to openai * test: fix rtc test failing due to change on transcript It also correctly setup database from sqlite, in case our configuration is setup to postgres. * ci: deactivate translation backend by default * test: fix modal->mock * refactor: implementing igor review, mock to passthrough	2025-08-04 12:07:30 -06:00
Mathieu Virbel	28ac031ff6	feat: use llamaindex everywhere (#525 ) * feat: use llamaindex for transcript final title too * refactor: removed llm backend, replaced with one single class+llamaindex * refactor: self-review * fix: typing * fix: tests * refactor: extract clean_title and add tests * test: fix * test: remove ensure_casing/nltk * fix: tiny mistake	2025-08-01 12:13:00 -06:00
Mathieu Virbel	ad56165b54	fix: remove unused settings and utils files (#522 ) * fix: remove unused settings and utils files * fix: remove migration done * fix: remove outdated scripts * fix: removing deployment of hermes, not used anymore * fix: partially remove secret, still have to understand frontend.	2025-07-31 17:45:48 -06:00
Mathieu Virbel	406164033d	feat: new summary using phi-4 and llama-index (#519 ) * feat: add litellm backend implementation * refactor: improve generate/completion methods for base LLM * refactor: remove tokenizer logic * style: apply code formatting * fix: remove hallucinations from LLM responses * refactor: comprehensive LLM and summarization rework * chore: remove debug code * feat: add structured output support to LiteLLM * refactor: apply self-review improvements * docs: add model structured output comments * docs: update model structured output comments * style: apply linting and formatting fixes * fix: resolve type logic bug * refactor: apply PR review feedback * refactor: apply additional PR review feedback * refactor: apply final PR review feedback * fix: improve schema passing for LLMs without structured output * feat: add PR comments and logger improvements * docs: update README and add HTTP logging * feat: improve HTTP logging * feat: add summary chunking functionality * fix: resolve title generation runtime issues * refactor: apply self-review improvements * style: apply linting and formatting * feat: implement LiteLLM class structure * style: apply linting and formatting fixes * docs: env template model name fix * chore: remove older litellm class * chore: format * refactor: simplify OpenAILLM * refactor: OpenAILLM tokenizer * refactor: self-review * refactor: self-review * refactor: self-review * chore: format * chore: remove LLM_USE_STRUCTURED_OUTPUT from envs * chore: roll back migration lint changes * chore: roll back migration lint changes * fix: make summary llm configuration optional for the tests * fix: missing f-string * fix: tweak the prompt for summary title * feat: try llamaindex for summarization * fix: complete refactor of summary builder using llamaindex and structured output when possible * fix: separate prompt as constant * fix: typings * fix: enhance prompt to prevent mentioning others subject while summarize one * fix: various changes after self-review * fix: from igor review --------- Co-authored-by: Igor Loskutov <igor.loskutoff@gmail.com>	2025-07-31 15:29:29 -06:00
Sergey Mankovsky	cfb1b2f9bc	Upgrade modal apps	2025-03-25 11:09:01 +01:00
Sergey Mankovsky	163d4a6e4a	Refactor transcribe segment	2025-01-20 12:46:20 +01:00
Sergey Mankovsky	99ff06ff17	OpenAI compatible transcription api	2025-01-20 12:27:58 +01:00
Sergey Mankovsky	7ff201f3ff	Fix model download	2024-12-27 14:23:03 +01:00
Mathieu Virbel	895ba36cb9	fix: modal upgrade (#421 )	2024-10-01 16:39:24 +02:00
Mathieu Virbel	5267ab2d37	feat: retake summary using NousResearch/Hermes-3-Llama-3.1-8B model (#415 ) This feature a new modal endpoint, and a complete new way to build the summary. ## SummaryBuilder The summary builder is based on conversational model, where an exchange between the model and the user is made. This allow more context inclusion and a better respect of the rules. It requires an endpoint with OpenAI-like completions endpoint (/v1/chat/completions) ## vLLM Hermes3 Unlike previous deployment, this one use vLLM, which gives OpenAI-like completions endpoint out of the box. It could also handle guided JSON generation, so jsonformer is not needed. But, the model is quite good to follow JSON schema if asked in the prompt. ## Conversion of long/short into summary builder The builder is identifying participants, find key subjects, get a summary for each, then get a quick recap. The quick recap is used as a short_summary, while the markdown including the quick recap + key subjects + summaries are used for the long_summary. This is why the nextjs component has to be updated, to correctly style h1 and keep the new line of the markdown.	2024-09-14 02:28:38 +02:00
Sara	004787c055	upgrade modal	2024-08-12 12:24:14 +02:00
projects-g	06b0abaf62	deployment fix (#364 )	2024-06-20 12:07:28 +05:30
projects-g	63502becd6	Move HF_token to modal secret (#354 ) * update all modal deployments and change seamless configuration due to change in src repo * add fixture * move token to secret	2024-04-19 10:30:45 +05:30
projects-g	72b22d1005	Update all modal deployments and change seamless configuration due to changes in src repo (#353 ) * update all modal deployments and change seamless configuration due to change in src repo * add fixture	2024-04-16 21:12:24 +05:30
Mathieu Virbel	8b1b71940f	hotfix/server: update diarization settings to increase timeout, reduce idle timeout on the minimum	2023-11-30 19:25:09 +01:00
projects-g	eae01c1495	Change diarization internal flow (#320 ) * change diarization internal flow	2023-11-30 22:00:06 +05:30
projects-g	5cb132cac7	fix loading shards from local cache (#313 )	2023-11-08 22:02:48 +05:30
Gokul Mohanarangan	894c989d60	update language codes	2023-10-14 17:35:30 +05:30
Sara	90c6824f52	replace two letter codes with three letter codes	2023-10-13 23:36:02 +02:00
Mathieu Virbel	9269db74c0	gpu: update format + list of country 2 to 3	2023-10-13 23:33:37 +02:00
Mathieu Virbel	6c1869b79a	gpu: improve concurrency on modal - coauthored with Gokul (#286 )	2023-10-13 21:15:57 +02:00
projects-g	1d92d43fe0	New summary (#283 ) * handover final summary to Zephyr deployment * fix display error * push new summary feature * fix failing test case * Added markdown support for final summary * update UI render issue * retain sentence tokenizer call --------- Co-authored-by: Koper <andreas@monadical.com>	2023-10-13 22:53:29 +05:30
projects-g	628c69f81c	Separate out transcription and translation into own Modal deployments (#268 ) * abstract transcript/translate into separate GPU apps * update app names * update transformers library version * update env.example file	2023-10-13 22:01:21 +05:30
Mathieu Virbel	47f7e1836e	server: remove warmup methods everywhere	2023-10-06 13:59:17 -04:00
projects-g	c9f613aff5	Revert GPU/Container retention settings for modal apps (#260 )	2023-10-03 09:54:32 +05:30
projects-g	6a43297309	Translation enhancements (#247 )	2023-09-26 19:49:54 +05:30
Gokul Mohanarangan	d7ed93ae3e	fix runtime download by creating specific storage paths for models	2023-09-25 09:34:42 +05:30
Gokul Mohanarangan	19dfb1d027	Upgrade to a bigger translation model	2023-09-20 20:02:52 +05:30
projects-g	9fe261406c	Feature additions (#210 ) * initial * add LLM features * update LLM logic * update llm functions: change control flow * add generation config * update return types * update processors and tests * update rtc_offer * revert new title processor change * fix unit tests * add comments and fix HTTP 500 * adjust prompt * test with reflector app * revert new event for final title * update * move onus onto processors * move onus onto processors * stash * add provision for gen config * dynamically pack the LLM input using context length * tune final summary params * update consolidated class structures * update consolidated class structures * update precommit * add broadcast processors * working baseline * Organize LLMParams * minor fixes * minor fixes * minor fixes * fix unit tests * fix unit tests * fix unit tests * update tests * update tests * edit pipeline response events * update summary return types * configure tests * alembic db migration * change LLM response flow * edit main llm functions * edit main llm functions * change llm name and gen cf * Update transcript_topic_detector.py * PR review comments * checkpoint before db event migration * update DB migration of past events * update DB migration of past events * edit LLM classes * Delete unwanted file * remove List typing * remove List typing * update oobabooga API call * topic enhancements * update UI event handling * move ensure_casing to llm base * update tests * update tests	2023-09-13 11:26:08 +05:30
Gokul Mohanarangan	9a7b89adaa	keep models in cache and load from cache	2023-09-08 10:05:17 +05:30
Gokul Mohanarangan	2bed312e64	persistent model storage	2023-09-08 00:22:38 +05:30
Gokul Mohanarangan	e613157fd6	update to use cache dir	2023-09-05 14:28:48 +05:30
Gokul Mohanarangan	6b84bbb4f6	download transcriber model	2023-09-05 12:52:07 +05:30
Gokul Mohanarangan	61e24969e4	change model download	2023-08-30 13:00:42 +05:30
Gokul Mohanarangan	012390d0aa	backup	2023-08-30 10:43:51 +05:30
Gokul Mohanarangan	e4fe3dfd3a	remove print	2023-08-28 15:28:53 +05:30
Gokul Mohanarangan	a5b8849e5e	change modal app name	2023-08-28 15:22:26 +05:30
Gokul Mohanarangan	d92a0de56c	update HTTP POST	2023-08-28 15:19:36 +05:30
Gokul Mohanarangan	ebbe01f282	update fixes	2023-08-28 14:32:21 +05:30
Gokul Mohanarangan	49d6e2d1dc	return both en and fr in transcriptio	2023-08-28 14:25:44 +05:30
Mathieu Virbel	d76bb83fe0	modal: fix schema passing issue with shadowing BaseModel.schema default	2023-08-22 17:10:36 +02:00
Gokul Mohanarangan	a0ea32db8a	review comments	2023-08-21 13:50:59 +05:30
Gokul Mohanarangan	acdd5f7dab	update	2023-08-21 12:53:49 +05:30
Gokul Mohanarangan	5b0883730f	translation update	2023-08-21 11:46:28 +05:30
Gokul Mohanarangan	2d686da15c	pass schema as dict	2023-08-17 21:51:44 +05:30
Gokul Mohanarangan	9103c8cca8	remove ast	2023-08-17 15:15:43 +05:30
Gokul Mohanarangan	2e48f89fdc	add comments and log	2023-08-17 09:33:59 +05:30
Gokul Mohanarangan	eb13a7bd64	make schema optional argument	2023-08-17 09:23:14 +05:30

1 2

53 Commits