Skip to content

StartUpdateAsync/GetResultAsync retries UpdateWorkflowExecution RPC internally at high frequency, ignoring RpcOptions.Retry=false #912

Description

@niketrehimes

Summary

WorkflowHandle.StartUpdateAsync / WorkflowUpdateHandle.GetResultAsync (and previously WorkflowHandle.ExecuteUpdateAsync) appear to retry the underlying UpdateWorkflowExecution gRPC call internally at a very high, backoff-less rate (measured: ~14 calls/second from a single StartUpdateAsync invocation) when the update's target lifecycle stage isn't reached quickly. Critically, RpcOptions.Retry = false has no effect on this — we tried it on StartUpdateAsync, GetResultAsync, and ExecuteUpdateAsync, in every combination, and the retry behavior was identical each time.

This causes sustained CPU usage (measured 40-70% of one core, for anywhere from ~30 seconds to several minutes per occurrence) in the calling .NET process, contributes to ResourceExhausted / "service rate limit exceeded" storms on the Temporal server side, and — because it appears to happen in native (Rust/tokio) code below the managed retry interceptor — makes the .NET host process unresponsive to graceful shutdown (Ctrl-C / SIGINT) while it's occurring.

Environment

  • Temporalio / Temporalio.Extensions.Hosting: reproduced on both 1.11.0 and 1.19.0
  • .NET SDK: 10.0.103, target framework net10.0
  • OS: macOS 26.6.2 (arm64)
  • Temporal server: temporalio/auto-setup:1.22.3 (single-container dev image, via Docker)
  • Deployment: local development machine, moderate concurrent load (Docker containers + multiple dotnet run processes running side by side)

What we observed

We have a Temporal Update handler ([WorkflowUpdate], returns a small result DTO) that a controller calls synchronously via ExecuteUpdateAsync (originally), later StartUpdateAsync(WaitForStage: Accepted) + GetResultAsync. The workflow itself always resolves the update's outcome (success or a thrown ApplicationFailureException) within a few seconds — this is corroborated by the Temporal server's own event history, which shows WorkflowExecutionUpdateCompleted typically 3-10 seconds after WorkflowExecutionUpdateAccepted.

Despite that, the .NET process's CPU stayed elevated for far longer than the workflow itself took to resolve the update — anywhere from ~30 seconds up to several minutes in our testing, well past the point where the server-side history already shows the update completed.

We captured a 6-second dotnet-trace (dotnet-sampled-thread-time profile) of the .NET process during one of these episodes. Aggregating by call frequency across all threads in that 6-second window:

84   Temporalio!Temporalio.Client.TemporalClient+Impl+<StartWorkflowUpdateAsync>d__44`1[System.__Canon].MoveNext()
84   Temporalio!Temporalio.Client.WorkflowService+Core.InvokeRpcAsync(...)
202  Temporalio!Temporalio.Bridge.Client+<CallAsync>d__15`1[System.__Canon].MoveNext()
168  Temporalio!Temporalio.Client.TemporalConnection+<InvokeRpcAsync>d__46`1[System.__Canon].MoveNext()
26   Temporalio!Temporalio.Api.WorkflowService.V1.UpdateWorkflowExecutionRequest ... (protobuf serialization)

That's 84 separate invocations of the internal StartWorkflowUpdateAsync/UpdateWorkflowExecution RPC path in 6 seconds, from a single await handle.StartUpdateAsync(...) call site in our code — i.e. the retrying is happening entirely inside the SDK's implementation, not from our own code calling it in a loop (we call it exactly once per user action).

A sample (macOS native profiler) taken during the same class of episode showed the native call stack repeatedly hitting temporalio_sdk_core_c_bridge::client::temporal_core_client_rpc_call → temporalio_client::retry::RetryClient<RC>::call → WorkflowService::update_workflow_execution, with UpdateWorkflowExecutionRequest/UpdateWorkflowExecutionResponse protobuf types appearing dozens of times in a 3-second sample window.

What we tried (none of it changed the behavior)

  • RpcOptions { Retry = false } on WorkflowUpdateOptions.Rpc (for ExecuteUpdateAsync)
  • RpcOptions { Retry = false, Timeout = TimeSpan.FromSeconds(8) } on WorkflowUpdateStartOptions.Rpc (for StartUpdateAsync)
  • RpcOptions { Retry = false, Timeout = TimeSpan.FromSeconds(8) } on the RpcOptions passed to GetResultAsync
  • All of the above combined
  • A CancellationToken with a bounded timeout on RpcOptions.CancellationToken
  • Upgrading from 1.11.0 → 1.19.0 (no change in behavior)

The only thing that reliably stops the elevated CPU is killing the process, or (we eventually did this) abandoning the Update API entirely in favor of a plain [WorkflowSignal] + our own application-level polling of side-effect state to learn the outcome.

Expected behavior

  • RpcOptions.Retry = false should suppress whatever internal retry/re-poll loop is driving these repeated UpdateWorkflowExecution calls, the same way it does for other high-level calls.
  • Failing that, some documented, application-controllable way to bound/backoff this specific retry loop would be very useful — today there doesn't appear to be one exposed above the native bridge.

Possible contributing factor

Our repro consistently involves a workflow that, in the course of handling the update, transitions through many workflow tasks in quick succession (a happy-path completion immediately followed by a failure in the next node's entry logic, causing a revert) under a resource-constrained host. It's possible the server is returning the update long-poll response without a terminal Outcome more often than usual under that load, and the client's internal retry-until-Completed loop has no backoff for that case specifically. Even so, we'd expect RpcOptions.Retry = false to have some observable effect, and it had none.

Happy to provide more detail, a minimal repro project, or the raw trace files if useful — this took a while to pin down and we have fairly complete diagnostic data (Temporal workflow event histories, dotnet-trace output, sample stacks) from several separate reproductions.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions