Summary
WorkflowHandle.StartUpdateAsync / WorkflowUpdateHandle.GetResultAsync (and previously WorkflowHandle.ExecuteUpdateAsync) appear to retry the underlying UpdateWorkflowExecution gRPC call internally at a very high, backoff-less rate (measured: ~14 calls/second from a single StartUpdateAsync invocation) when the update's target lifecycle stage isn't reached quickly. Critically, RpcOptions.Retry = false has no effect on this — we tried it on StartUpdateAsync, GetResultAsync, and ExecuteUpdateAsync, in every combination, and the retry behavior was identical each time.
This causes sustained CPU usage (measured 40-70% of one core, for anywhere from ~30 seconds to several minutes per occurrence) in the calling .NET process, contributes to ResourceExhausted / "service rate limit exceeded" storms on the Temporal server side, and — because it appears to happen in native (Rust/tokio) code below the managed retry interceptor — makes the .NET host process unresponsive to graceful shutdown (Ctrl-C / SIGINT) while it's occurring.
Environment
Temporalio / Temporalio.Extensions.Hosting: reproduced on both 1.11.0 and 1.19.0
- .NET SDK: 10.0.103, target framework
net10.0
- OS: macOS 26.6.2 (arm64)
- Temporal server:
temporalio/auto-setup:1.22.3 (single-container dev image, via Docker)
- Deployment: local development machine, moderate concurrent load (Docker containers + multiple
dotnet run processes running side by side)
What we observed
We have a Temporal Update handler ([WorkflowUpdate], returns a small result DTO) that a controller calls synchronously via ExecuteUpdateAsync (originally), later StartUpdateAsync(WaitForStage: Accepted) + GetResultAsync. The workflow itself always resolves the update's outcome (success or a thrown ApplicationFailureException) within a few seconds — this is corroborated by the Temporal server's own event history, which shows WorkflowExecutionUpdateCompleted typically 3-10 seconds after WorkflowExecutionUpdateAccepted.
Despite that, the .NET process's CPU stayed elevated for far longer than the workflow itself took to resolve the update — anywhere from ~30 seconds up to several minutes in our testing, well past the point where the server-side history already shows the update completed.
We captured a 6-second dotnet-trace (dotnet-sampled-thread-time profile) of the .NET process during one of these episodes. Aggregating by call frequency across all threads in that 6-second window:
84 Temporalio!Temporalio.Client.TemporalClient+Impl+<StartWorkflowUpdateAsync>d__44`1[System.__Canon].MoveNext()
84 Temporalio!Temporalio.Client.WorkflowService+Core.InvokeRpcAsync(...)
202 Temporalio!Temporalio.Bridge.Client+<CallAsync>d__15`1[System.__Canon].MoveNext()
168 Temporalio!Temporalio.Client.TemporalConnection+<InvokeRpcAsync>d__46`1[System.__Canon].MoveNext()
26 Temporalio!Temporalio.Api.WorkflowService.V1.UpdateWorkflowExecutionRequest ... (protobuf serialization)
That's 84 separate invocations of the internal StartWorkflowUpdateAsync/UpdateWorkflowExecution RPC path in 6 seconds, from a single await handle.StartUpdateAsync(...) call site in our code — i.e. the retrying is happening entirely inside the SDK's implementation, not from our own code calling it in a loop (we call it exactly once per user action).
A sample (macOS native profiler) taken during the same class of episode showed the native call stack repeatedly hitting temporalio_sdk_core_c_bridge::client::temporal_core_client_rpc_call → temporalio_client::retry::RetryClient<RC>::call → WorkflowService::update_workflow_execution, with UpdateWorkflowExecutionRequest/UpdateWorkflowExecutionResponse protobuf types appearing dozens of times in a 3-second sample window.
What we tried (none of it changed the behavior)
RpcOptions { Retry = false } on WorkflowUpdateOptions.Rpc (for ExecuteUpdateAsync)
RpcOptions { Retry = false, Timeout = TimeSpan.FromSeconds(8) } on WorkflowUpdateStartOptions.Rpc (for StartUpdateAsync)
RpcOptions { Retry = false, Timeout = TimeSpan.FromSeconds(8) } on the RpcOptions passed to GetResultAsync
- All of the above combined
- A
CancellationToken with a bounded timeout on RpcOptions.CancellationToken
- Upgrading from 1.11.0 → 1.19.0 (no change in behavior)
The only thing that reliably stops the elevated CPU is killing the process, or (we eventually did this) abandoning the Update API entirely in favor of a plain [WorkflowSignal] + our own application-level polling of side-effect state to learn the outcome.
Expected behavior
RpcOptions.Retry = false should suppress whatever internal retry/re-poll loop is driving these repeated UpdateWorkflowExecution calls, the same way it does for other high-level calls.
- Failing that, some documented, application-controllable way to bound/backoff this specific retry loop would be very useful — today there doesn't appear to be one exposed above the native bridge.
Possible contributing factor
Our repro consistently involves a workflow that, in the course of handling the update, transitions through many workflow tasks in quick succession (a happy-path completion immediately followed by a failure in the next node's entry logic, causing a revert) under a resource-constrained host. It's possible the server is returning the update long-poll response without a terminal Outcome more often than usual under that load, and the client's internal retry-until-Completed loop has no backoff for that case specifically. Even so, we'd expect RpcOptions.Retry = false to have some observable effect, and it had none.
Happy to provide more detail, a minimal repro project, or the raw trace files if useful — this took a while to pin down and we have fairly complete diagnostic data (Temporal workflow event histories, dotnet-trace output, sample stacks) from several separate reproductions.
Summary
WorkflowHandle.StartUpdateAsync/WorkflowUpdateHandle.GetResultAsync(and previouslyWorkflowHandle.ExecuteUpdateAsync) appear to retry the underlyingUpdateWorkflowExecutiongRPC call internally at a very high, backoff-less rate (measured: ~14 calls/second from a singleStartUpdateAsyncinvocation) when the update's target lifecycle stage isn't reached quickly. Critically,RpcOptions.Retry = falsehas no effect on this — we tried it onStartUpdateAsync,GetResultAsync, andExecuteUpdateAsync, in every combination, and the retry behavior was identical each time.This causes sustained CPU usage (measured 40-70% of one core, for anywhere from ~30 seconds to several minutes per occurrence) in the calling .NET process, contributes to
ResourceExhausted/ "service rate limit exceeded" storms on the Temporal server side, and — because it appears to happen in native (Rust/tokio) code below the managed retry interceptor — makes the .NET host process unresponsive to graceful shutdown (Ctrl-C / SIGINT) while it's occurring.Environment
Temporalio/Temporalio.Extensions.Hosting: reproduced on both 1.11.0 and 1.19.0net10.0temporalio/auto-setup:1.22.3(single-container dev image, via Docker)dotnet runprocesses running side by side)What we observed
We have a Temporal Update handler (
[WorkflowUpdate], returns a small result DTO) that a controller calls synchronously viaExecuteUpdateAsync(originally), laterStartUpdateAsync(WaitForStage: Accepted)+GetResultAsync. The workflow itself always resolves the update's outcome (success or a thrownApplicationFailureException) within a few seconds — this is corroborated by the Temporal server's own event history, which showsWorkflowExecutionUpdateCompletedtypically 3-10 seconds afterWorkflowExecutionUpdateAccepted.Despite that, the .NET process's CPU stayed elevated for far longer than the workflow itself took to resolve the update — anywhere from ~30 seconds up to several minutes in our testing, well past the point where the server-side history already shows the update completed.
We captured a 6-second
dotnet-trace(dotnet-sampled-thread-timeprofile) of the .NET process during one of these episodes. Aggregating by call frequency across all threads in that 6-second window:That's 84 separate invocations of the internal
StartWorkflowUpdateAsync/UpdateWorkflowExecutionRPC path in 6 seconds, from a singleawait handle.StartUpdateAsync(...)call site in our code — i.e. the retrying is happening entirely inside the SDK's implementation, not from our own code calling it in a loop (we call it exactly once per user action).A
sample(macOS native profiler) taken during the same class of episode showed the native call stack repeatedly hittingtemporalio_sdk_core_c_bridge::client::temporal_core_client_rpc_call→temporalio_client::retry::RetryClient<RC>::call→WorkflowService::update_workflow_execution, withUpdateWorkflowExecutionRequest/UpdateWorkflowExecutionResponseprotobuf types appearing dozens of times in a 3-second sample window.What we tried (none of it changed the behavior)
RpcOptions { Retry = false }onWorkflowUpdateOptions.Rpc(forExecuteUpdateAsync)RpcOptions { Retry = false, Timeout = TimeSpan.FromSeconds(8) }onWorkflowUpdateStartOptions.Rpc(forStartUpdateAsync)RpcOptions { Retry = false, Timeout = TimeSpan.FromSeconds(8) }on theRpcOptionspassed toGetResultAsyncCancellationTokenwith a bounded timeout onRpcOptions.CancellationTokenThe only thing that reliably stops the elevated CPU is killing the process, or (we eventually did this) abandoning the Update API entirely in favor of a plain
[WorkflowSignal]+ our own application-level polling of side-effect state to learn the outcome.Expected behavior
RpcOptions.Retry = falseshould suppress whatever internal retry/re-poll loop is driving these repeatedUpdateWorkflowExecutioncalls, the same way it does for other high-level calls.Possible contributing factor
Our repro consistently involves a workflow that, in the course of handling the update, transitions through many workflow tasks in quick succession (a happy-path completion immediately followed by a failure in the next node's entry logic, causing a revert) under a resource-constrained host. It's possible the server is returning the update long-poll response without a terminal
Outcomemore often than usual under that load, and the client's internal retry-until-Completedloop has no backoff for that case specifically. Even so, we'd expectRpcOptions.Retry = falseto have some observable effect, and it had none.Happy to provide more detail, a minimal repro project, or the raw trace files if useful — this took a while to pin down and we have fairly complete diagnostic data (Temporal workflow event histories,
dotnet-traceoutput,samplestacks) from several separate reproductions.