Files
open-webui-ai4me/backend/open_webui/routers
Classic298andMichael efe5416f83 fix: reduce TTFT by caching model lookups in chat completion (#20886)
fix: reduce TTFT by caching model lookups in chat completion

Skip expensive get_all_models() calls when models are already cached
in app.state. This significantly reduces Time To First Token (TTFT)
for chat completions and embeddings requests.

Previously, every request called get_all_models() which fetches model
lists from all configured backends. Now we check the cache first and
only call get_all_models() on cache miss.

Affected endpoints:
- openai: generate_chat_completion, embeddings
- ollama: embed, embeddings

Fixes #20069

Co-authored-by: Michael <42099345+mickeytheseal@users.noreply.github.com>
2026-02-11 18:29:10 -06:00
..
2026-02-11 16:24:11 -06:00
2026-02-11 17:56:49 -06:00
2026-02-11 18:25:37 -06:00
2026-02-11 16:24:11 -06:00
2026-01-29 18:51:02 +04:00
2026-01-22 03:11:33 +04:00
2026-01-09 20:44:31 +04:00
2026-02-09 13:28:14 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 15:55:23 -06:00
2026-02-10 15:41:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 18:24:30 -06:00
2026-02-11 16:24:11 -06:00
2026-02-06 22:33:49 +04:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00
2026-02-11 16:24:11 -06:00