Skip to content

Controller marks Tasks Failed permanently when Substrate is briefly unavailable #421

Description

@ghchinoy

Summary

If the Substrate Control API is unreachable when ax-controller reconciles a Task, the Task is set to phase: Failed and the event is acknowledged. Nothing reconciles it again after Substrate recovers, so the Task stays Failed until the user deletes and re-applies it. Delete events have the same problem: the Task stays in Terminating until the user runs ax delete again.

A Substrate restart or rolling upgrade is enough to trigger this.

Related, but not covering this case: #367 (the same behavior for ResourceExhausted), #351 (unacknowledged events are not recovered), and #418 (events acknowledged when the status write fails).

Reproduce

AX d0bc38b, ax-server and ax-controller with Redis, and a Substrate Control endpoint that is not yet listening:

  1. Start ax-controller with --substrate-endpoint pointing at an endpoint that is not yet ready.
  2. ax apply -f examples/simple.yaml
  3. Start the Substrate endpoint.
  4. ax describe task simple-task still shows:
Phase:        Failed
Conditions:
  Ready  False  AtespaceCreationFailed  creating atespace "default": rpc error: code = Unavailable desc = connection error: desc = "transport: Error while dialing: dial tcp 127.0.0.1:50051: connect: connection refused"
  1. ax delete task simple-task right after the endpoint comes back:
level=ERROR msg="error processing task event" action=delete error="cleaning up task default/simple-task: deleting actor default/simple-task: rpc error: code = Unavailable desc = connection error: ... connect: connection refused"

The Task stays in Terminating. Running ax delete again a minute later succeeds.

In step 5, the endpoint had been up for about 20 seconds when the delete failed. The controller's gRPC connection was still in reconnect backoff from the earlier failures. internal/substrate/client.go uses gRPC defaults, and the backoff grows to about 2 minutes.

Cause

  • internal/controller/worker.go:92-103: Worker.Run logs the error from processEvent and then always calls sub.Ack.
  • internal/controller/reconciler.go:121-122, 187-188 and 210-211 set Phase = "Failed" for any error, including codes.Unavailable.
  • No periodic resync reconciles Tasks that are not Running.

Suggested fix

Any of these, or a combination:

  1. Treat codes.Unavailable (and DeadlineExceeded) as transient: keep the Task Pending with a condition like Ready=False, reason=SubstrateUnavailable, and don't acknowledge the event, so it's retried. This depends on Redis consumer does not recover unacknowledged events after worker failure #351's pending-entry recovery, or on re-publishing the event with a delay.
  2. Use grpc.WaitForReady(true) with a bounded per-call timeout, so a call made during reconnect waits for the connection instead of failing immediately.
  3. Add a periodic resync that re-reconciles Tasks in Pending, Failed or Terminating.

Option 1 matches the fix proposed in #367 for ResourceExhausted, so both could share one "transient Substrate error" path.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions