reflector

mirror of https://github.com/Monadical-SAS/reflector.git synced 2026-02-05 02:16:46 +00:00

Author	SHA1	Message	Date
Mathieu Virbel	9eab952c63	feat: postgresql migration and removal of sqlite in pytest (#546 ) * feat: remove support of sqlite, 100% postgres * fix: more migration and make datetime timezone aware in postgres * fix: change how database is get, and use contextvar to have difference instance between different loops * test: properly use client fixture that handle lifetime/database connection * fix: add missing client fixture parameters to test functions This commit fixes NameError issues where test functions were trying to use the 'client' fixture but didn't have it as a parameter. The changes include: 1. Added 'client' parameter to test functions in: - test_transcripts_audio_download.py (6 functions including fixture) - test_transcripts_speaker.py (3 functions) - test_transcripts_upload.py (1 function) - test_transcripts_rtc_ws.py (2 functions + appserver fixture) 2. Resolved naming conflicts in test_transcripts_rtc_ws.py where both HTTP client and StreamClient were using variable name 'client'. StreamClient instances are now named 'stream_client' to avoid conflicts. 3. Added missing 'from reflector.app import app' import in rtc_ws tests. Background: Previously implemented contextvars solution with get_database() function resolves asyncio event loop conflicts in Celery tasks. The global client fixture was also created to replace manual AsyncClient instances, ensuring proper FastAPI application lifecycle management and database connections during tests. All tests now pass except for 2 pre-existing RTC WebSocket test failures related to asyncpg connection issues unrelated to these fixes. * fix: ensure task are correctly closed * fix: make separate event loop for the live server * fix: make default settings pointing at postgres * build: remove pytest-docker deps out of dev, just tests group	2025-08-14 11:40:52 -06:00
Igor Loskutov	6fb5cb21c2	feat: search backend (#537 ) * docs: transient docs * chore: cleanup * webvtt WIP * webvtt field * chore: webvtt tests comments * chore: remove useless tests * feat: search TASK.md * feat: full text search by title/webvtt * chore: search api task * feat: search api * feat: search API * chore: rm task md * chore: roll back unnecessary validators * chore: pr review WIP * chore: pr review WIP * chore: pr review * chore: top imports * feat: better lint + ci * feat: better lint + ci * feat: better lint + ci * feat: better lint + ci * chore: lint * chore: lint * fix: db datetime definitions * fix: flush() params * fix: update transcript mutability expectation / test * fix: update transcript mutability expectation / test * chore: auto review * chore: new controller extraction * chore: new controller extraction * chore: cleanup * chore: review WIP * chore: pr WIP * chore: remove ci lint * chore: openapi regeneration * chore: openapi regeneration * chore: postgres test doc * fix: .dockerignore for arm binaries * fix: .dockerignore for arm binaries * fix: cap test loops * fix: cap test loops * fix: cap test loops * fix: get_transcript_topics * chore: remove flow.md docs and claude guidance * chore: remove claude.md db doc * chore: remove claude.md db doc * chore: remove claude.md db doc * chore: remove claude.md db doc	2025-08-13 10:03:38 -04:00
Igor Loskutov	a42ed12982	fix: evaluation cli event wrap (#536 ) * fix: evaluation cli event wrap * fix: evaluation cli event wrap * chore: remove unrelated change * chore: rollback claude.md changes	2025-08-11 19:28:52 -04:00
Mathieu Virbel	dc177af3ff	feat: implement service-specific Modal API keys with auto processor pattern (#528 ) * fix: refactor modal API key configuration for better separation of concerns - Split generic MODAL_API_KEY into service-specific keys: - TRANSCRIPT_API_KEY for transcription service - DIARIZATION_API_KEY for diarization service - TRANSLATE_API_KEY for translation service - Remove deprecated _MODAL_API_KEY settings - Add proper validation to ensure URLs are set when using modal processors - Update README with new configuration format BREAKING CHANGE: Configuration keys have changed. Update your .env file: - TRANSCRIPT_MODAL_API_KEY → TRANSCRIPT_API_KEY - LLM_MODAL_API_KEY → (removed, use TRANSCRIPT_API_KEY) - Add DIARIZATION_API_KEY and TRANSLATE_API_KEY if using those services fix: update Modal backend configuration to use service-specific API keys - Changed from generic MODAL_API_KEY to service-specific keys: - TRANSCRIPT_MODAL_API_KEY for transcription - DIARIZATION_MODAL_API_KEY for diarization - TRANSLATION_MODAL_API_KEY for translation - Updated audio_transcript_modal.py and audio_diarization_modal.py to use modal_api_key parameter - Updated documentation in README.md, CLAUDE.md, and env.example * feat: implement auto/modal pattern for translation processor - Created TranscriptTranslatorAutoProcessor following the same pattern as transcript/diarization - Created TranscriptTranslatorModalProcessor with TRANSLATION_MODAL_API_KEY support - Added TRANSLATION_BACKEND setting (defaults to "modal") - Updated all imports to use TranscriptTranslatorAutoProcessor instead of TranscriptTranslatorProcessor - Updated env.example with TRANSLATION_BACKEND and TRANSLATION_MODAL_API_KEY - Updated test to expect TranscriptTranslatorModalProcessor name - All tests passing * refactor: simplify transcript_translator base class to match other processors - Moved all implementation from base class to modal processor - Base class now only defines abstract _translate method - Follows the same minimal pattern as audio_diarization and audio_transcript base classes - Updated test mock to use _translate instead of get_translation - All tests passing * chore: clean up settings and improve type annotations - Remove deprecated generic API key variables from settings - Add comments to group Modal-specific settings - Improve type annotations for modal_api_key parameters * fix: typing * fix: passing key to openai * test: fix rtc test failing due to change on transcript It also correctly setup database from sqlite, in case our configuration is setup to postgres. * ci: deactivate translation backend by default * test: fix modal->mock * refactor: implementing igor review, mock to passthrough	2025-08-04 12:07:30 -06:00
Mathieu Virbel	28ac031ff6	feat: use llamaindex everywhere (#525 ) * feat: use llamaindex for transcript final title too * refactor: removed llm backend, replaced with one single class+llamaindex * refactor: self-review * fix: typing * fix: tests * refactor: extract clean_title and add tests * test: fix * test: remove ensure_casing/nltk * fix: tiny mistake	2025-08-01 12:13:00 -06:00
Mathieu Virbel	f5b82d44e3	style: use ruff for linting and formatting (#524 )	2025-07-31 17:57:43 -06:00
Mathieu Virbel	406164033d	feat: new summary using phi-4 and llama-index (#519 ) * feat: add litellm backend implementation * refactor: improve generate/completion methods for base LLM * refactor: remove tokenizer logic * style: apply code formatting * fix: remove hallucinations from LLM responses * refactor: comprehensive LLM and summarization rework * chore: remove debug code * feat: add structured output support to LiteLLM * refactor: apply self-review improvements * docs: add model structured output comments * docs: update model structured output comments * style: apply linting and formatting fixes * fix: resolve type logic bug * refactor: apply PR review feedback * refactor: apply additional PR review feedback * refactor: apply final PR review feedback * fix: improve schema passing for LLMs without structured output * feat: add PR comments and logger improvements * docs: update README and add HTTP logging * feat: improve HTTP logging * feat: add summary chunking functionality * fix: resolve title generation runtime issues * refactor: apply self-review improvements * style: apply linting and formatting * feat: implement LiteLLM class structure * style: apply linting and formatting fixes * docs: env template model name fix * chore: remove older litellm class * chore: format * refactor: simplify OpenAILLM * refactor: OpenAILLM tokenizer * refactor: self-review * refactor: self-review * refactor: self-review * chore: format * chore: remove LLM_USE_STRUCTURED_OUTPUT from envs * chore: roll back migration lint changes * chore: roll back migration lint changes * fix: make summary llm configuration optional for the tests * fix: missing f-string * fix: tweak the prompt for summary title * feat: try llamaindex for summarization * fix: complete refactor of summary builder using llamaindex and structured output when possible * fix: separate prompt as constant * fix: typings * fix: enhance prompt to prevent mentioning others subject while summarize one * fix: various changes after self-review * fix: from igor review --------- Co-authored-by: Igor Loskutov <igor.loskutoff@gmail.com>	2025-07-31 15:29:29 -06:00
Sergey Mankovsky	43562391b7	Fix duration	2025-01-24 14:47:43 +01:00
Sergey Mankovsky	b601f18d2d	Fix summary generation	2025-01-21 16:53:09 +01:00
Sergey Mankovsky	49be4013bc	Remove unused headers	2025-01-20 12:50:53 +01:00
Sergey Mankovsky	99ff06ff17	OpenAI compatible transcription api	2025-01-20 12:27:58 +01:00
Mathieu Virbel	5267ab2d37	feat: retake summary using NousResearch/Hermes-3-Llama-3.1-8B model (#415 ) This feature a new modal endpoint, and a complete new way to build the summary. ## SummaryBuilder The summary builder is based on conversational model, where an exchange between the model and the user is made. This allow more context inclusion and a better respect of the rules. It requires an endpoint with OpenAI-like completions endpoint (/v1/chat/completions) ## vLLM Hermes3 Unlike previous deployment, this one use vLLM, which gives OpenAI-like completions endpoint out of the box. It could also handle guided JSON generation, so jsonformer is not needed. But, the model is quite good to follow JSON schema if asked in the prompt. ## Conversion of long/short into summary builder The builder is identifying participants, find key subjects, get a summary for each, then get a quick recap. The quick recap is used as a short_summary, while the markdown including the quick recap + key subjects + summaries are used for the long_summary. This is why the nextjs component has to be updated, to correctly style h1 and keep the new line of the markdown.	2024-09-14 02:28:38 +02:00
Mathieu Virbel	07b29d42a7	server: add topic duration, and endpoint for getting words group per speaker on a topic	2023-12-11 19:46:05 +01:00
Mathieu Virbel	f9ca92a15c	Merge pull request #328 from Monadical-SAS/feat-enhance-diarization Enhance diarization results	2023-11-30 17:31:50 +01:00
projects-g	eae01c1495	Change diarization internal flow (#320 ) * change diarization internal flow	2023-11-30 22:00:06 +05:30
Mathieu Virbel	3ebb21923b	server: enhance diarization algorithm	2023-11-29 20:34:43 +01:00
Sara	3ba764ba86	Merge branch 'main' of github.com:Monadical-SAS/reflector into sara/loading-states	2023-11-17 15:33:47 +01:00
Sara	a846e38fbd	fix waveform in pipeline	2023-11-17 13:38:32 +01:00
Sara	1fc261a669	try to move waveform to pipeline	2023-11-15 20:30:00 +01:00
Mathieu Virbel	4dbcb80228	server: remove reference to banana.dev	2023-11-15 19:49:32 +01:00
Mathieu Virbel	84e425bd3b	server: fix slow profanity filter	2023-11-11 01:00:28 +01:00
Mathieu Virbel	c255b41475	server: fix crash when translator is stopped without having a single push	2023-11-11 01:00:09 +01:00
Mathieu Virbel	e18a7c8d4e	server: correctly save duration, when filewriter is finished	2023-11-11 01:00:09 +01:00
Mathieu Virbel	9642d0fd1e	hotfix/server: fix duplication of topics	2023-11-02 19:40:45 +01:00
Mathieu Virbel	c87c30d339	hotfix/server: add follow_redirect on modal	2023-11-02 19:09:13 +01:00
Mathieu Virbel	057c636c56	server: move logging to base implementation, not specialization	2023-11-02 17:39:21 +01:00
Mathieu Virbel	19b5ba2c4c	server: add diarization logger information	2023-11-02 17:39:21 +01:00
Mathieu Virbel	4da890b95f	server: add dummy diarization and fixes instanciation	2023-11-02 17:39:21 +01:00
Mathieu Virbel	d8a842f099	server: full diarization processor implementation based on gokul app	2023-11-02 17:39:21 +01:00
Mathieu Virbel	07c4d080c2	server: refactor with diarization, logic works	2023-11-02 17:39:21 +01:00
Mathieu Virbel	1c42473da0	server: refactor with clearer pipeline instanciation and linked to model	2023-11-02 17:39:21 +01:00
Mathieu Virbel	367912869d	server: make processors in broadcast to be executed in parallel	2023-11-02 17:39:21 +01:00
Mathieu Virbel	f4cffc0e66	server: add tests on segmentation and fix issue with speaker	2023-11-02 17:39:21 +01:00
Mathieu Virbel	00eb9bbf3c	server: improve split algorithm	2023-11-02 17:39:21 +01:00
Mathieu Virbel	b323254376	server: move out profanity filter to transcript, and implement segmentation	2023-11-02 17:39:21 +01:00
Gokul Mohanarangan	c1a9005ec3	update buller condition	2023-10-14 18:55:40 +05:30
Gokul Mohanarangan	79fa537c35	update return format	2023-10-14 18:08:16 +05:30
Gokul Mohanarangan	894c989d60	update language codes	2023-10-14 17:35:30 +05:30
Sara	90c6824f52	replace two letter codes with three letter codes	2023-10-13 23:36:02 +02:00
projects-g	1d92d43fe0	New summary (#283 ) * handover final summary to Zephyr deployment * fix display error * push new summary feature * fix failing test case * Added markdown support for final summary * update UI render issue * retain sentence tokenizer call --------- Co-authored-by: Koper <andreas@monadical.com>	2023-10-13 22:53:29 +05:30
projects-g	628c69f81c	Separate out transcription and translation into own Modal deployments (#268 ) * abstract transcript/translate into separate GPU apps * update app names * update transformers library version * update env.example file	2023-10-13 22:01:21 +05:30
Mathieu Virbel	47f7e1836e	server: remove warmup methods everywhere	2023-10-06 13:59:17 -04:00
projects-g	e78bcc9190	Scaleai Translation (#258 ) * hotfix * remove assert from translation * review comments * reflector.media change targetLang to en	2023-09-28 18:16:39 +05:30
projects-g	24aa9a74bd	hotfix (#254 )	2023-09-27 19:20:43 +05:30
projects-g	6a43297309	Translation enhancements (#247 )	2023-09-26 19:49:54 +05:30
Gokul Mohanarangan	f56eaeb6cc	dont delete censored words	2023-09-25 21:25:18 +05:30
Gokul Mohanarangan	80fd5e6176	update llm params	2023-09-22 07:49:41 +05:30
Gokul Mohanarangan	009d52ea23	update casing and trimming	2023-09-22 07:29:01 +05:30
Gokul Mohanarangan	ab41ce90e8	add profanity filter, post-process topic/title	2023-09-21 11:12:00 +05:30
Mathieu Virbel	2b9eef6131	server: use mp3 as default for audio storage Closes #223	2023-09-13 17:26:03 +02:00

1 2

88 Commits