Mahmoud Ashraf and GitHub
d57c5b40b0
Remove the usage of transformers.pipeline from BatchedInferencePipeline and fix word timestamps for batched inference ( #921 )
...
* fix word timestamps for batched inference
* remove hf pipeline
2024-07-27 09:02:58 +07:00
zh-plus and GitHub
83a368e98a
Make vad-related parameters configurable for batched inference. ( #923 )
2024-07-24 09:00:32 +07:00
eb8390233c
New PR for Faster Whisper: Batching Support, Speed Boosts, and Quality Enhancements ( #856 )
...
Batching Support, Speed Boosts, and Quality Enhancements
---------
Co-authored-by: Hargun Mujral <83234565+hargunmujral@users.noreply.github.com >
Co-authored-by: MahmoudAshraf97 <hassouna97.ma@gmail.com >
2024-07-18 16:48:52 +07:00
trungkienbkhn and GitHub
fbcf58bf98
Fix language detection with non-speech audio ( #895 )
2024-07-05 14:43:45 +07:00
Jordi Mas and GitHub
1195359984
Filter out non_speech_tokens in suppressed tokens ( #898 )
...
* Filter out non_speech_tokens in suppressed tokens
2024-07-05 14:43:11 +07:00
trungkienbkhn and GitHub
c22db5125d
Bump version to 1.0.3 ( #887 )
2024-07-01 16:36:12 +07:00
ABen and GitHub
8862bee1f8
Improve language detection when using clip_timestamps ( #867 )
2024-07-01 16:12:45 +07:00
Ki Hoon Kim and GitHub
8d400e9870
Upgrade to Silero-Vad V5 ( #884 )
...
* Fix window_size_samples to 512
* Update SileroVADModel
* Replace ONNX file with V5 version
2024-07-01 15:40:37 +07:00
Napuh and GitHub
f53be1e811
Add distil models to WhisperModel init and download_model docstrings ( #847 )
...
* chore: add distil models to WhisperModel init docstring and download_model docstring
2024-05-20 08:51:22 +07:00
Natanael Tan and GitHub
4acdb5c619
Fix #839 incorrect clip_timestamps being used in model ( #842 )
...
* Fix #839
Changed the code from updating the TranscriptionOptions class instead of the options object which likely was the cause of unexpected behaviour
2024-05-17 16:35:07 +07:00
trungkienbkhn and GitHub
2f6913efc8
Bump version to 1.0.2 ( #816 )
2024-05-06 09:02:54 +07:00
Keating Reid and GitHub
49a80eb8a8
Clarify documentation for hotwords ( #817 )
...
* Clarify documentation for hotwords
* Remove redundant type specifications
2024-05-06 08:52:59 +07:00
trungkienbkhn and GitHub
8d5e6d56d9
Support initializing more whisper model args ( #807 )
2024-05-04 15:12:59 +07:00
847fec4492
Feature/add hotwords ( #731 )
...
* add hotword params
---------
Co-authored-by: jax <jax_builder@gamil.com >
2024-05-04 15:11:52 +07:00
otakutyrant and GitHub
91c8307aa6
make faster_whisper.assets as a valid python package to distribute ( #772 ) ( #774 )
2024-04-02 18:22:22 +02:00
Purfview and GitHub
b024972a56
Foolproof: Disable VAD if clip_timestamps is in use ( #769 )
...
* Foolproof: Disable VAD if clip_timestamps is in use
Prevent silly things to happen.
2024-04-02 18:20:34 +02:00
Purfview and GitHub
8ae82c8372
Bugfix: code breaks if audio is empty ( #768 )
...
* Bugfix: code breaks if audio is empty
Regression since https://github.com/SYSTRAN/faster-whisper/pull/732 PR
2024-04-02 18:18:12 +02:00
trungkienbkhn and GitHub
e0c3a9ed34
Update project github link to SYSTRAN ( #746 )
2024-03-27 08:31:17 +01:00
Sanchit Gandhi and GitHub
a67e0e47ae
Add support for distil-large-v3 ( #755 )
...
* add distil-large-v3
* Update README.md
* use fp16 weights from Systran
2024-03-26 14:58:39 +01:00
trungkienbkhn and GitHub
1eb9a8004c
Improve language detection ( #732 )
2024-03-12 15:44:49 +01:00
trungkienbkhn and GitHub
a342b028b7
Bump version to 1.0.1 ( #725 )
2024-03-01 11:32:12 +01:00
Purfview and GitHub
5090cc9d0d
Fix window end heuristic for hallucination_silence_threshold ( #706 )
...
Removes the wishful heuristic causing more issues than it's fixing.
Same as https://github.com/openai/whisper/pull/2043
Example of the issue: https://github.com/openai/whisper/pull/1838#issuecomment-1960041500
2024-02-29 17:59:32 +01:00
trungkienbkhn and GitHub
16141e65d9
Add pad_or_trim function to handle segment before encoding ( #705 )
2024-02-29 17:08:28 +01:00
trungkienbkhn and GitHub
06d32bf0c1
Bump version to 1.0.0 ( #696 )
2024-02-22 09:49:01 +01:00
Purfview and GitHub
30d6043e90
Prevent infinite loop for out-of-bound timestamps in clip_timestamps ( #697 )
...
Same as https://github.com/openai/whisper/pull/2005
2024-02-22 09:48:35 +01:00
trungkienbkhn and GitHub
092067208b
Add clip_timestamps and hallucination_silence_threshold options ( #646 )
2024-02-20 17:34:54 +01:00
Purfview and GitHub
3aec421849
Add: More clarity of what "max_new_tokens" does ( #658 )
...
* Add: More clarity of what "max_new_tokens" does
2024-01-28 21:40:33 +01:00
Purfview and GitHub
00efce1696
Bugfix: Illogical "Avoid computing higher temperatures on no_speech" ( #652 )
2024-01-24 11:54:43 +01:00
metame and GitHub
ad3c83045b
support distil-whisper ( #557 )
2024-01-24 10:17:12 +01:00
Purfview and GitHub
ebcfd6b964
Fix broken prompt_reset_on_temperature ( #604 )
...
* Fix broken prompt_reset_on_temperature
Fixing: https://github.com/SYSTRAN/faster-whisper/issues/603
Broken because `generate_with_fallback()` doesn't return final temperature.
Regression since PR356 -> https://github.com/SYSTRAN/faster-whisper/pull/356
2023-12-13 13:14:39 +01:00
trungkienbkhn and GitHub
19329a3611
Word timing tweaks ( #616 )
2023-12-13 12:38:44 +01:00
Clayton Yochum and GitHub
9641d5f56a
Force read-mode in av.open ( #566 )
...
The `av.open` functions checks input metadata to determine the mode to open with ("r" or "w"). If an input to `decode_audio` is found to be in write-mode, without this change it can't be read. Forcing read mode solves this.
2023-11-27 10:43:35 +01:00
Dang Chuan Nguyen
e1a218fab1
Bump version to 0.10.0
2023-11-24 23:19:47 +01:00
3084409633
Add V3 Support ( #578 )
...
* Add V3 Support
* update conversion example
---------
Co-authored-by: oscaarjs <oscar.johansson@conversy.se >
2023-11-24 23:16:12 +01:00
Guillaume Klein
5a0541ea7d
Bump version to 0.9.0
2023-09-18 16:21:37 +02:00
Guillaume Klein and GitHub
e94711bb5c
Add property WhisperModel.supported_languages ( #476 )
...
* Expose function supported_languages
* Make it a method
2023-09-14 17:42:02 +02:00
Guillaume Klein and GitHub
0048844f54
Expose function available_models ( #475 )
...
* Expose function available_models
* Add test case
2023-09-14 17:17:01 +02:00
Guillaume Klein
a49097e655
Add some missing typing annotations in transcribe.py
2023-09-12 15:45:54 +02:00
Guillaume Klein and GitHub
81086f6d33
Always run the encoder at the beginning of the loop ( #468 )
2023-09-12 14:44:37 +02:00
Guillaume Klein and GitHub
727ab81f31
Improve error message for invalid task and language parameters ( #466 )
2023-09-12 10:02:23 +02:00
Guillaume Klein
ad388cd394
Bump version to 0.8.0
2023-09-04 11:56:48 +02:00
Guillaume Klein and GitHub
4a41746e55
Log a warning when the model is English-only but the language is set to something else ( #454 )
2023-09-04 11:55:40 +02:00
Guillaume Klein and GitHub
1e6eb967c9
Add "large" alias for "large-v2" model ( #453 )
2023-09-04 11:54:42 +02:00
Guillaume Klein and GitHub
f0ff12965a
Expose generation parameter no_repeat_ngram_size ( #449 )
2023-09-01 17:31:30 +02:00
Guillaume Klein and GitHub
5871858a5f
Force the garbage collector to run after decoding the audio with PyAV ( #448 )
2023-09-01 15:25:13 +02:00
MinorJinx and GitHub
e87fbf8a49
Added audio duration after VAD to TranscriptionInfo object ( #445 )
...
* Added VAD removed audio duration to TranscriptionInfo object
Along with the duration of the original audio, this commit adds the seconds of audio removed by the VAD to the returned info obj
* Chaning naming for duration_after_vad
Instead of the property returning the audio duration removed, it now returns the final duration after the vad.
If vad_filter is False or if it doesn't remove any audio, the original duration is returned.
2023-08-31 17:19:48 +02:00
1562b02345
added repetition_penalty to TranscriptionOptions ( #403 )
...
Co-authored-by: Aisu Wata <aisu.wata0@gmail.com >
2023-08-06 10:08:24 +02:00
Purfview and GitHub
1ce16652ee
Adds DEBUG log message for prompt_reset_on_temperature ( #399 )
...
Produce DEBUG log message if prompt_reset_on_temperature threshold is met.
2023-08-04 09:06:17 +02:00
Purfview and GitHub
857be6f621
Rename clear_previous_text_on_temperature argument ( #398 )
...
`prompt_reset_on_temperature` is more clear what it does.
2023-08-03 18:44:37 +02:00
KH and GitHub
1a1eb1a027
Add clear_previous_text_on_temperature parameter ( #397 )
...
* Add clear_previous_text_on_temperature parameter
* Add a description
2023-08-03 15:40:58 +02:00