Skip to content

LCORE-4070: raise a model error when OGX reports a streamed request as failed - #2868

Open
max-svistunov wants to merge 1 commit into
lightspeed-core:mainfrom
max-svistunov:lcore-4070-streamed-failure-error
Open

max-svistunov wants to merge 1 commit into
lightspeed-core:mainfrom
max-svistunov:lcore-4070-streamed-failure-error

Conversation

@max-svistunov

@max-svistunov max-svistunov commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Description

LCORE-4070. When a query is too long for the model's context window, /v1/streaming_query streamed a complete answer that had nothing to do with the question, where Prompt is too long was expected. The change comes from the LCORE-4108 work, whose compaction capability has a stream hook and so streams the model call on /v1/query too. It is proposed alone: it fixes LCORE-4070 on main and needs nothing from LCORE-4108.

Cause. OGX reports a provider failure of a streamed Responses request inside the stream, as a response.failed event. OgxResponsesModel passed the event on, and pydantic-ai took the empty response for a missing answer: it sent a second model request with its retry prompt (Validation feedback: Please return text. Fix the errors and try again.). With a conversation attached, _prepare_conversation_continuation keeps only the messages after the failed response, so that request carried the retry prompt and not the question. Its reply was streamed as the answer and stored as the turn.

Fix. _FilteredResponseStream raises ModelHTTPError for the event (status 500, the error of the event as body), so the run ends after one model request. The endpoints already catch agent errors and map them: a context-length message gives 413 Prompt is too long, a message that names RESOURCE_EXHAUSTED gives 429, anything else the generic 500. A provider rate limit is in the last group: the event carries a code and a message, no HTTP status, and the code is not read.

What changes. Only requests whose model call is streamed and which OGX reports as failed:

Path Before After
/v1/streaming_query token events and turn_complete with the reply to the retry prompt; that turn is stored one error event; nothing stored
/v1/streaming_query, failure after part of the answer (by reading, not run) turn_complete with the partial text, then error; that turn is stored the tokens, then the mapped error; nothing stored
A2A (by reading, not run) the reply to the retry prompt as the result task failed with the mapped message
/v1/query with a Granite Guardian shield, whose stream hook makes the run stream (by reading, not run) 200 with the reply to the retry prompt the mapped HTTP error

What does not change. /v1/query without Granite Guardian uses the non-streamed call. In library mode it answers 503 to the same query, before and after: the OGX exception escapes the library transport there, a separate defect outside this PR. A request that succeeds is handled as before.

What is given up. Until now a failed streamed request was followed by a second one, so a failure that happened once (a rate limit, a provider 5xx) could still end in a 200 stream. That second request did not repeat the question, so what came back was the reply to the retry prompt, stored as the turn. (On a compacted turn it carried the query, without the summaries and recent turns; by reading.) Now the client gets the error and has to resend the request. This PR adds no retry of its own.

Type of change

  • Refactor
  • New feature
  • Bug fix
  • CVE fix
  • Optimization
  • Documentation Update
  • Configuration Update
  • Bump-up service version
  • Bump-up dependent library [pyproject.toml + uv.lock]
  • Bump-up dependent library [requirements.*.txt for Konflux]
  • Bump-up library or tool used for development (does not change the final image)
  • CI configuration change
  • Konflux configuration change
  • Unit tests improvement
  • Integration tests improvement
  • End to end tests improvement
  • Benchmarks improvement

Tools used to create PR

Identify any AI code assistants used in this PR (for transparency and review context)

  • Assisted-by: Claude Opus 4.8
  • Generated by: Claude Opus 4.8

Related Tickets & Documents

  • Related Issue # LCORE-4070
  • Closes # LCORE-4070

Checklist before requesting a review

  • I have performed a self-review of my code.
  • PR has passed all pre-merge test jobs.
  • If it is a core feature, I have added thorough tests.

Testing

  1. Reproduce the bug on a running service, then check it is fixed. Start the service in library mode with the default configuration and openai/gpt-4o-mini, and put a recording proxy between OGX and the provider. Send a query far beyond the context window ("the quick brown fox jumps over the lazy dog " * 17000 + "Reply with OK only.", 748,019 characters) to /v1/streaming_query. Expected: one provider request, an error event with status 413, nothing stored. Without the fix:
data: {"event": "start", "data": {"conversation_id": "6b0103551d2c...", "request_id": "8b395262-..."}}
data: {"event": "token", "data": {"id": 0, "token": "It"}}          ... and 57 more token events
data: {"event": "turn_complete", "data": {"id": 58, "token": "It seems like you received feedback indicating that there are errors that need to be corrected in your text. ..."}}
provider: 400 context_length_exceeded   user: <748019 chars> the quick brown fox jumps over the lazy dog ...
provider: 200                           user: Validation feedback: Please return text.  Fix the errors and try again.
provider: 400 context_length_exceeded   topic summary of the same query
stored:   user 'Validation feedback:\nPlease return text.\n\nFix the errors and try again.', assistant 'It seems like you received feedback ...'

The stream ends there, with no end event: in library mode the failed topic-summary request raises through the endpoint. The ticket reports a generic error event at that point. With the fix (one provider request, rejected with the same 400; the conversation holds no item):

data: {"event": "start", "data": {"conversation_id": "ee674bb23598...", "request_id": "ca2b1539-..."}}
data: {"event": "error", "data": {"status_code": 413, "response": "Prompt is too long", "cause": "The input exceeds the context window size of model 'openai/gpt-4o-mini'."}}
  1. Check that nothing else moved. Send the same query to /v1/query (non-streamed call): with and without the fix it answers 503 {"detail":{"response":"Unable to connect to OGX","cause":"Connection error while trying to reach backend service."}} after three provider requests. Send What is the capital of France? Reply with one word. to /v1/streaming_query: both streams are start, two token events, turn_complete with Paris. and end, with two provider requests each. Server mode, A2A and Granite Guardian were not run.

  2. Run the tests for this change, the full suites and the linters. Both new tests fail without the change in src/ (DID NOT RAISE ModelHTTPError; UnexpectedModelBehavior: Exceeded maximum output retries (1)):

uv run pytest tests/unit/pydantic_ai_lightspeed/ogx/test_model.py -q                # 51 passed
uv run pytest tests/unit -q                                                        # 3709 passed, 1 skipped
uv run pytest tests/integration --ignore=tests/integration/container_lifecycle -q   # 326 passed
uv run make format                                                                  # no changes
uv run make black ruff docstyle pylint pyright                                      # all pass

Summary by CodeRabbit

  • Bug Fixes
    • Streamed failures from the OGX provider now surface as model errors instead of being silently missed.
    • Context-length failures during an agent run are now reported as prompt-too-long errors, with the request made only once.

@coderabbitai

coderabbitai Bot commented Oct 7, 2026 •

Copy link
Copy Markdown
Contributor

Review in Change Stack →

No actionable comments were generated in the recent review. 🎉

ℹ️ Recent review info
⚙️ Run configuration
  • Configuration used: Repository: lightspeed-core/lightspeed-stack/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 4275e98e-a86b-4085-a9a6-75f8af6c4ed4
📥 Commits

Reviewing files that changed from the base of the PR and between d8eb7cf and 99c9d27.

📒 Files selected for processing (2)
  • src/pydantic_ai_lightspeed/ogx/_model.py
  • tests/unit/pydantic_ai_lightspeed/ogx/test_model.py

Included review availability: This review used your included allowance. Your plan provides up to 1 included review per hour; 0 remain after this review.

📜 Recent review details
⏰ Context from checks skipped due to timeout. (35)
  • GitHub Check: E2E: server / ci / default
  • GitHub Check: E2E: server / ci / rbac
  • GitHub Check: E2E: library / ci / mcp
  • GitHub Check: E2E: server / ci / authorized
  • GitHub Check: E2E: server / ci / shields
  • GitHub Check: E2E: library / ci / authorized
  • GitHub Check: E2E: server / ci / mcp
  • GitHub Check: E2E: server / ci / other
  • GitHub Check: E2E: library / ci / shields
  • GitHub Check: E2E: library / ci / skills
  • GitHub Check: E2E: server / ci / tls
  • GitHub Check: E2E: library / ci / other
  • GitHub Check: E2E: library / ci / rbac
  • GitHub Check: E2E: library / ci / default
  • GitHub Check: E2E: server / ci / skills
  • GitHub Check: ruff
  • GitHub Check: integration_tests (3.12)
  • GitHub Check: Pyright
  • GitHub Check: bandit
  • GitHub Check: radon
  • GitHub Check: unit_tests (3.13)
  • GitHub Check: mypy
  • GitHub Check: pydocstyle
  • GitHub Check: unit_tests (3.12)
  • GitHub Check: integration_tests (3.13)
  • GitHub Check: list_outdated_dependencies
  • GitHub Check: build-pr
  • GitHub Check: black
  • GitHub Check: Red Hat Konflux / rag-content-0-8-e2e-tests / lightspeed-stack-0-8
  • GitHub Check: spectral
  • GitHub Check: shellcheck
  • GitHub Check: Pylinter
  • GitHub Check: Red Hat Konflux / lightspeed-core-0-8-enterprise-contract / lightspeed-stack-0-8
  • GitHub Check: Red Hat Konflux / lightspeed-stack-0-8-e2e-tests / lightspeed-stack-0-8
  • GitHub Check: Konflux kflux-prd-rh02 / lightspeed-stack-0-8-on-pull-request
🧰 Additional context used
🪛 ast-grep (0.45.3)
tests/unit/pydantic_ai_lightspeed/ogx/test_model.py

[warning] 739-739: Configuring an LLM/agent client endpoint over http:// sends prompts and responses (and often API keys) in cleartext, exposing them to interception. Use https for the base_url.
Context: base_url="http://localhost:8321/v1"
Note: [CWE-319] Cleartext Transmission of Sensitive Information.

(llm-client-insecure-http-python)

🔇 Additional comments (2)
src/pydantic_ai_lightspeed/ogx/_model.py (1)

34-34: LGTM!

Also applies to: 154-158, 164-174

tests/unit/pydantic_ai_lightspeed/ogx/test_model.py (1)

3-3: LGTM!

Also applies to: 14-15, 25-25, 27-27, 34-34, 44-49, 563-576, 697-723, 725-765


Walkthrough

The OGX stream handler now raises ModelHTTPError when it receives a failed response event. Tests cover the HTTP status, the agent run’s request count, and mapping a context-length error to PromptTooLongResponse.

Changes

OGX streamed failures

Layer / File(s) Summary
Handle failed response events
src/pydantic_ai_lightspeed/ogx/_model.py, tests/unit/pydantic_ai_lightspeed/ogx/test_model.py
The stream handler raises ModelHTTPError with status 500, the response model name, and the serialized error body when available. Tests verify the raised error and confirm that the agent run makes one request and maps a context-length error to PromptTooLongResponse.

Priority: ➖ Normal

Estimated code review effort: 2 (Simple) | ~10 minutes

Change: Bug fix

Suggested reviewers: asimurka

Merge Risk: ⚪ Minimal · up to 99c9d

Streamed model failures are handled through the existing error responses. No issue identified here prevents merging after normal checks.

🚥 Pre-merge checks | ✅ 7
✅ Passed checks (7 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly and concisely describes the main change: raising a model error when OGX reports a streamed request failure.
Docstring Coverage ✅ Passed Docstring coverage is 100.00% which is sufficient. The required threshold is 80.00%. Docstring coverage is scoped to functions touched by this diff. Analyzed 7 functions across 2 files.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
Performance And Algorithmic Complexity ✅ Passed No meaningful performance regression is introduced. The PR adds one isinstance check per streamed event and calls error.model_dump() only once when a terminal ResponseFailedEvent occurs. It adds…
Security And Secret Handling ✅ Passed PASSED. The PR changes only src/pydantic_ai_lightspeed/ogx/_model.py and its unit test. The new code raises ModelHTTPError with the OGX failure body at lines 164-174; existing agent handlers map t…
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create a new PR
✨ Simplify code
  • Create a new PR
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

…s failed

A prompt too long for the model made /v1/streaming_query stream a
complete answer that had nothing to do with the question, where
"Prompt is too long" was expected. The conversation then held a retry
prompt as the user message.

OGX reports a provider failure of a streamed Responses request inside
the stream, as a response.failed event. OgxResponsesModel passed it on,
pydantic-ai took the empty response for a missing answer and sent a
second request with the retry prompt "Validation feedback: Please
return text. Fix the errors and try again." Its reply was streamed and
stored as the answer.

_FilteredResponseStream now raises ModelHTTPError for the event, with
status 500 and the error of the event as body. The run ends after one
model request and nothing is stored. The existing mapping turns a
context-length message into 413; an unrecognised failure stays a 500.

This applies where the model request is streamed: /v1/streaming_query,
A2A, and /v1/query with a Granite Guardian shield. A failed request is
no longer followed by a second one. The non-streamed call is untouched.

Two unit tests: request_stream raises for a response.failed event, and
an agent run over such a stream makes one model request and maps to
413. Checked on a live service in library mode with gpt-4o-mini on
/v1/streaming_query only; A2A and Granite Guardian were not run.
@max-svistunov
max-svistunov force-pushed the lcore-4070-streamed-failure-error branch from 24f281d to 99c9d27 Compare October 8, 2026 16:00

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant