-
Type:
New Feature
-
Resolution: Unresolved
-
None
-
Affects Version/s: None
-
Component/s: translate5 AI
-
Medium
-
None
-
None
-
Emptyshow more show less
Problem
The OpenAI/translate5AI plugin keeps every model's capabilities in a hardcoded PHP constant - MODEL_CAPABILITIES in PrivatePlugins/OpenAI/Api/Config.php:77. Each entry defines the model's max context window, max output tokens, reasoning flag, allowed reasoning-effort values and unsupported chat params, matched to a model name by str_contains in a manually ordered, most-specific-first list. We add support for new models frequently. When a new model is not in this list (or doesn't match a pattern), the code silently falls back to conservative defaults (MIN_OUTPUT_TOKENS = 8192, MIN_FRAME_SIZE = 32768 - Config.php:230/:236). That under-utilises the model - batches are sized far smaller than the model could actually handle - and the only way to fix it is a code change and a release. There is no way for an admin to register or correct a model's capabilities from the UI.
Two further problems sit on top of this:
- Batch sizing is a blunt fixed count. How many segments go into one request is governed by maxBatchSize (per resource, default 6 - ModelProperties.php:131), on top of a token check (Connector.php:487 isAllowedByContentSize). A single fixed count can't adapt to how much the model can actually take, which depends on the model's context/output limits and the size of the prompt (which the user can inflate with RAG content - ConversationFactory.php:128). A value that is safe for one model/prompt wastes capacity on another.
- The user can't see what will actually be sent. Because the prompt size is variable (RAG/system messages are user-configurable per resource), the number of segments that fit in a batch is not obvious. Today it is invisible until pre-translation runs.
Solution
Bring the capabilities registry into an admin UI, make batch sizing adaptive, and show the user the resulting segment count.
- Global capabilities registry (DB + admin grid). Move MODEL_CAPABILITIES into a DB table (LEK_openai_model_capabilities), seeded by migration from the current constant (kept in code as the seed + ultimate fallback). Add an admin-only CRUD grid to edit rows (model pattern, max context, max output, reasoning, reasoning-effort options, unsupported params). The matching rule changes from manual ordering to longest-matching-pattern wins, so grid row order can't break resolution. Config's existing static accessors stay as the seam (delegating to a ModelCapabilitiesRepository), so all call sites are untouched. Unknown models keep today's safe fallback.
- Dynamic batch size: low / medium / high presets. Replace the fixed per-resource maxBatchSize count with a preset. A preset is a fraction of the model's effective usable budget - min(context-window room after the prompt, output-token room / response-size ratio) - so it respects the output-token ceiling that actually binds first (with the default 5x response ratio, the output limit is reached long before the context window). Keep a hidden "manual" count for power users, and keep useBatchedPretranslation (on/off). The runtime batching consumes the computed token budget instead of a hard count.
- Live segments-per-batch preview. Compute the estimate with a small service (BatchSizeCalculator) that mirrors the runtime token math, using the real assembled prompt (incl. RAG, tokenized with the existing Gpt3Tokenizer). Show a rough estimate in the resource settings dialog (approx context size + approx segments/batch for the chosen preset; no task context), and the exact figure as a column in the per-task language-resource grid once the resource is assigned to a task (real prompt + that task's segment sizes). Disclose it as an estimate (per-segment terminology varies) and warn when a large prompt shrinks the budget to ~1 segment.