Skip to content

spark.comet.shuffle.jvm.batchSize=0 hangs a task, and an unknown spark.comet.exec.memoryPool fails every task #6259

Description

@andygrove

Describe the bug

Two settings accept values that fail only once a task runs on an executor.

spark.comet.shuffle.jvm.batchSize is only checked to be no larger than spark.comet.batchSize (CometConf.scala#L677-L688), so 0 is accepted. process_sorted_row_partition then never advances, because n = min(batch_size, ...) is 0 on every pass of its loop (row.rs#L1395-L1396). The loop runs inside a JNI call, so killing the task doesn't stop it.

spark.comet.exec.memoryPool has no checkValues (CometConf.scala#L905-L913). A misspelled value, including a difference in case only, gets through, and in off-heap mode every task then fails with Unsupported memory pool type when it creates a native plan.

Expected behavior

Both are validated when they are set. The batch size must be positive, and the pool type must be fair_unified or greedy_unified, accepted in any case and lowercased before it reaches native code.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

area:memoryMemory pools, reservations, OOM handlingarea:shuffleShuffle (JVM and native)bugSomething isn't workingrequires-triage

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions