You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
If the Substrate Control API is unreachable when ax-controller reconciles a Task, the Task is set to phase: Failed and the event is acknowledged. Nothing reconciles it again after Substrate recovers, so the Task stays Failed until the user deletes and re-applies it. Delete events have the same problem: the Task stays in Terminating until the user runs ax delete again.
A Substrate restart or rolling upgrade is enough to trigger this.
Related, but not covering this case: #367 (the same behavior for ResourceExhausted), #351 (unacknowledged events are not recovered), and #418 (events acknowledged when the status write fails).
Reproduce
AX d0bc38b, ax-server and ax-controller with Redis, and a Substrate Control endpoint that is not yet listening:
Start ax-controller with --substrate-endpoint pointing at an endpoint that is not yet ready.
The Task stays in Terminating. Running ax delete again a minute later succeeds.
In step 5, the endpoint had been up for about 20 seconds when the delete failed. The controller's gRPC connection was still in reconnect backoff from the earlier failures. internal/substrate/client.go uses gRPC defaults, and the backoff grows to about 2 minutes.
Cause
internal/controller/worker.go:92-103: Worker.Run logs the error from processEvent and then always calls sub.Ack.
internal/controller/reconciler.go:121-122, 187-188 and 210-211 set Phase = "Failed" for any error, including codes.Unavailable.
No periodic resync reconciles Tasks that are not Running.
Suggested fix
Any of these, or a combination:
Treat codes.Unavailable (and DeadlineExceeded) as transient: keep the Task Pending with a condition like Ready=False, reason=SubstrateUnavailable, and don't acknowledge the event, so it's retried. This depends on Redis consumer does not recover unacknowledged events after worker failure #351's pending-entry recovery, or on re-publishing the event with a delay.
Use grpc.WaitForReady(true) with a bounded per-call timeout, so a call made during reconnect waits for the connection instead of failing immediately.
Add a periodic resync that re-reconciles Tasks in Pending, Failed or Terminating.
Option 1 matches the fix proposed in #367 for ResourceExhausted, so both could share one "transient Substrate error" path.
Summary
If the Substrate Control API is unreachable when
ax-controllerreconciles a Task, the Task is set tophase: Failedand the event is acknowledged. Nothing reconciles it again after Substrate recovers, so the Task staysFaileduntil the user deletes and re-applies it. Delete events have the same problem: the Task stays inTerminatinguntil the user runsax deleteagain.A Substrate restart or rolling upgrade is enough to trigger this.
Related, but not covering this case: #367 (the same behavior for
ResourceExhausted), #351 (unacknowledged events are not recovered), and #418 (events acknowledged when the status write fails).Reproduce
AX
d0bc38b,ax-serverandax-controllerwith Redis, and a Substrate Control endpoint that is not yet listening:ax-controllerwith--substrate-endpointpointing at an endpoint that is not yet ready.ax apply -f examples/simple.yamlax describe task simple-taskstill shows:ax delete task simple-taskright after the endpoint comes back:The Task stays in
Terminating. Runningax deleteagain a minute later succeeds.In step 5, the endpoint had been up for about 20 seconds when the delete failed. The controller's gRPC connection was still in reconnect backoff from the earlier failures.
internal/substrate/client.gouses gRPC defaults, and the backoff grows to about 2 minutes.Cause
internal/controller/worker.go:92-103:Worker.Runlogs the error fromprocessEventand then always callssub.Ack.internal/controller/reconciler.go:121-122,187-188and210-211setPhase = "Failed"for any error, includingcodes.Unavailable.Running.Suggested fix
Any of these, or a combination:
codes.Unavailable(andDeadlineExceeded) as transient: keep the TaskPendingwith a condition likeReady=False, reason=SubstrateUnavailable, and don't acknowledge the event, so it's retried. This depends on Redis consumer does not recover unacknowledged events after worker failure #351's pending-entry recovery, or on re-publishing the event with a delay.grpc.WaitForReady(true)with a bounded per-call timeout, so a call made during reconnect waits for the connection instead of failing immediately.Pending,FailedorTerminating.Option 1 matches the fix proposed in #367 for
ResourceExhausted, so both could share one "transient Substrate error" path.