Translate5AI: request priority parameter for self hosted models

XMLWordPrintable

    • Type: Task
    • Resolution: Unresolved
    • None
    • Affects Version/s: None
    • Component/s: translate5 AI

      translate5AI will be used in production by an external party. Their requests must always be processed first. Requests coming from translate5 (testing/internal) should run with lower priority and may wait when the server is busy with production traffic.

      To enable this, the translate5AI server supports a priority parameter on each request. It controls the processing order: a lower number means higher priority, and requests without the parameter get the default value 0 (highest priority). When the server is under load, low-priority requests are queued or paused so high-priority requests are handled immediately.

      Part 1 – translate5AI server:

      Activate priority-based scheduling in the server configuration and restart the service. After this, the server sorts incoming requests by the priority value instead of plain first-come-first-served.

      Part 2 – translate5:

      Add the parameter priority with value 100 to every request body sent from translate5 to translate5AI:
       

      $client->chat()->create([
          'model'    => 'mistral-small-4',
          'messages' => [...],
          'priority' => 100,
      ]); 

      Production traffic sends no priority parameter, so it automatically gets 0 and is always served first. Add the parameter in one central place (HTTP client wrapper) so it applies to all translate5 requests.

            Assignee:
            Aleksandar Mitrev
            Reporter:
            Aleksandar Mitrev
            None
            None
            Leon Kiz
            Stephan Bergmann, Sylvia Schumacher
            Votes:
            0 Vote for this issue
            Watchers:
            1 Start watching this issue

              Created:
              Updated:
              None
              None