Skip to main content

DynamoEval Workers

DynamoEval runs its attacks in KEDA ScaledJob workers that scale from zero. When a test never starts, or starts and fails immediately, the platform itself usually reports healthy: the HelmRelease is Ready, the API and UI are Running, and nothing in the namespace looks wrong. Each section below is keyed to the literal symptom, so match what you are seeing and skip the rest.

Worker Pods Are Short-Lived by Design

A worker takes one job, finishes it, and exits. The Job is then deleted after ttlSecondsAfterFinished, which defaults to 60 seconds, so the pod disappears before most people can read its logs. To keep a pod around while debugging, raise the value on the job you are investigating:

Retain finished worker pods for an hour
ttlSecondsAfterFinished: 3600
failedJobsHistoryLimit: 10

Lower it again afterwards, because retained pods hold node capacity.

Synthetic Data Generation Rejects the Configured Provider​

A System Policy Compliance or Policy Jailbreak test fails within seconds:

guardrail compliant prompt generation failed for all generation models (<provider>/<model>).
Errors: <provider>/<model>: ValueError: Model <provider>/<model> not supported

This is not the provider rejecting the model, and the credentials are not wrong. The message is produced inside the worker before any network call. Synthetic data generation dispatches on the provider prefix of the model ID, and raises for any prefix it does not recognise. Every model under an unrecognised prefix produces the identical error.

PrefixRecognised for generation
together_ai/Yes
mistral/Yes
openai/Yes
gemini/Yes
anthropic/Yes
bedrock/Yes
azure/, azure_ai/Yes, from platform release 3.26.7
Anything else, including vertex_ai/No

Two attack types generate synthetic prompts before attacking, so only those two are affected:

AttackUI nameAffected
guardrail_benchmarkSystem Policy ComplianceYes
policy_jailbreakPolicy JailbreakYes
static_jailbreakStatic JailbreakNo
adaptive_jailbreakAdaptive JailbreakNo

Generation is a separate path from judging. Judges, attacker models, and the target system route through a component that supports a wider provider set, which is why the rest of a deployment works while only these two attacks fail.

Fix​

Point generation at a provider the worker recognises and the cluster can reach. Leaving the generation models unset is not a fix on its own: the defaults are anthropic/claude-haiku-4-5, openai/gpt-4.1-mini, and gemini/gemini-2.5-flash, which an isolated network cannot reach.

Three environment variables set the generation models, one per fallback slot. The worker tries them in order until one succeeds:

  • DATA_GENERATION_ANTHROPIC_MODEL_ID
  • DATA_GENERATION_OPENAI_MODEL_ID
  • DATA_GENERATION_GEMINI_MODEL_ID

The variable names reflect the default provider in each slot, not a restriction: any recognised prefix is valid in any of the three.

Set all three slots to a Bedrock model ID:

Route generation to Bedrock
DATA_GENERATION_ANTHROPIC_MODEL_ID=bedrock/<model-id>
DATA_GENERATION_OPENAI_MODEL_ID=bedrock/<model-id>
DATA_GENERATION_GEMINI_MODEL_ID=bedrock/<model-id>

Bedrock generation reads four AWS variables, and all four are mandatory, including AWS_SESSION_TOKEN:

VariableRequired
AWS_ACCESS_KEY_IDYes
AWS_SECRET_ACCESS_KEYYes
AWS_REGION_NAMEYes
AWS_SESSION_TOKENYes

Generation reads these variables directly, so it does not pick up credentials from an instance role or from IAM Roles for Service Accounts. A missing variable fails the worker at configuration time rather than at the API call.

Setting DYNAMOEVAL_JUDGE_MODE to AWS_BEDROCK routes judges and attacker models to Bedrock, but it does not redirect generation. The three variables above are still required.

Two model-name formats coexist, and mixing them up produces errors that read as though the model itself is wrong:

VariableFormatExample
AZURE_MODEL_NAMEBare deployment name<deployment>
SYNDATA_MODEL_NAME, DATA_GENERATION_*_MODEL_IDProvider-prefixedazure/<deployment>

A bare name where a prefix is expected fails with LLM Provider NOT provided.

Workers Keep Running the Previous Build After an Image Tag Change​

The ScaledJob shows the new tag, the HelmRelease reports Ready, and worker behaviour is unchanged, including errors that the new build is supposed to fix.

Check the full image reference, not just the tag:

Read the image reference a worker actually uses
kubectl -n <namespace> get scaledjob <scaledjob-name> \
-o jsonpath='{.spec.jobTargetRef.template.spec.containers[0].image}{"\n"}'

If the output ends in @sha256:..., the values set both a tag and a digest. The chart concatenates them into repository:tag@digest, and container runtimes resolve by digest and ignore the tag entirely. Updating only the tag changes what the object displays and nothing about what runs.

Fix​

Either remove the digest value so the tag is authoritative, or update the digest in the same change as the tag. Read the digest for the new tag from the registry:

Get the digest for a tag
curl -sI -u "<username>:<password>" \
-H "Accept: application/vnd.docker.distribution.manifest.v2+json" \
"https://<registry-host>/v2/<project>/images/pen-testing/manifests/<tag>" \
| grep -i docker-content-digest
pullPolicy Is Not the Key the Chart Reads

The ScaledJob template reads image.imagePullPolicy. A value set as image.pullPolicy is silently ignored, so the pull policy falls back to the default. This has no effect while a digest is pinned, because digests are content-addressed, but it matters once the digest is removed.

Image Pull Fails With context canceled​

Worker pods appear stuck for several minutes and then fail:

Failed to pull image "<registry>/<project>/images/pen-testing:<tag>":
rpc error: code = Canceled desc = failed to pull and unpack image
"<registry>/<project>/images/pen-testing:<tag>": context canceled

Error: ErrImagePull
Back-off pulling image "<registry>/<project>/images/pen-testing:<tag>"

Canceled is distinct from not found and unauthorized. The image exists and the credentials work; the transfer was aborted part-way. Kubelet cancels a pull that stops making measurable progress rather than waiting indefinitely.

Worker images are large, and one layer dominates the transfer. Two conditions commonly combine to stall it:

  • The registry is a proxying remote and the tag is a cache miss. The proxy fetches the layer from upstream before serving it downstream, during which the node observes no progress.
  • Several workers pull the same cold tag at once. KEDA creates a pod per queued job regardless of whether earlier pods have started, so bandwidth splits and every transfer slows toward the cancellation threshold.

Fix​

Warm the registry cache with one deliberate pull from a machine that tolerates a long transfer, then let the workers pull from cache:

Warm the cache before workers try
skopeo copy \
docker://<registry>/<project>/images/pen-testing:<tag> \
dir:/tmp/discard

A docker pull against the same reference works equally well. While warming, reduce maxReplicaCount on the affected job so the cluster is not attempting several cold pulls in parallel, then restore it.

Confirm the node can hold the extracted image if pulls still fail once the cache is warm:

Check for disk pressure on the target node
kubectl describe node <node-name> | grep -A 6 Conditions

Sizing guidance and the pre-pull pattern that removes this failure mode are on the Container Image Delivery page.

Tests Stay at Preparing Resources With No Worker Pods​

Preparing Resources means the attack is registered and every stage is still at 0%, so no worker has begun executing. Work down this chain, because each step has a different owner.

DynamoEval workers do not read the queue through KEDA's Redis scaler. They use the external scaler, so the path is:

API -> pushes job to Redis -> metrics-server reads queue length
-> reports over gRPC to KEDA -> KEDA scales the ScaledJob -> worker pod runs

A break anywhere in that chain looks identical from the UI.

Step 1: is there a worker pod at all?
kubectl -n <namespace> get pods | grep scaled-job

A pod in Pending phase is not necessarily unscheduled. A pod stays Pending until a container starts, which includes the whole image pull. Read its events before concluding anything:

Step 2: what is the pod waiting on?
kubectl -n <namespace> describe pod <pod-name> | tail -25

PodScheduled: True followed by Pulling means the pod has a node and is downloading. FailedScheduling names the constraint instead, commonly insufficient ephemeral-storage, because each worker requests a fixed reservation on top of a multi-gigabyte image.

Step 3: if no pod exists, is KEDA scaling?
kubectl -n <namespace> get scaledjobs

READY=True with ACTIVE=False and zero pods is the correct idle state, not a fault. ACTIVE=False while a test is genuinely queued means the scaler is reporting a queue length of zero; check that metrics-server is running and that it reads the same Redis database the API writes to. A mismatched database reports zero indefinitely.

READY=False means KEDA cannot reach the scaler at all. Verify the Service and its endpoints:

Confirm the scaler endpoint resolves
kubectl -n <namespace> get svc,endpoints dynamoai-metrics-server
Step 4: is the KEDA operator itself healthy?
kubectl -n keda get pods
kubectl -n keda logs deploy/keda-operator --tail=50 | grep -E 'external_scaler|Scaling Jobs'

Restarts on keda-operator, or repeated external_scaler errors reporting context canceled, indicate the operator is terminating mid-reconcile. It cannot then read queue lengths or write status conditions, so nothing scales. Identify the cause before restarting, because the restart clears the evidence:

Why the operator terminated
kubectl -n keda describe pod <keda-operator-pod> | grep -A 15 "Last State"
kubectl -n keda logs <keda-operator-pod> --previous --tail=50

OOMKilled calls for a higher memory limit. A non-zero exit with lease-renewal messages in the previous logs points at leader election against a slow API server.

Worker Pods Are Recreated After a Test Is Cancelled​

The test no longer appears in the UI, but KEDA keeps creating worker pods for its queue, often several, each replaced as it exits.

Cancelling a test removes its message from Redis only for attacks the API still finds in an IN_PROGRESS or QUEUED state, matched by job ID. If the attack record is already gone or in another state, cancellation returns NO_ACTION and the queued message is left behind. The scaler still sees a non-empty queue and keeps requesting workers, and no worker can complete the job because the test it belongs to no longer exists.

Fix​

Stop the churn first, then clear the message.

1. Pause the affected ScaledJob
kubectl -n <namespace> annotate scaledjob <scaledjob-name> \
autoscaling.keda.sh/paused="true" --overwrite

Confirm it took effect. The Paused condition reports ScaledJobPaused:

2. Confirm the pause
kubectl -n <namespace> get scaledjob <scaledjob-name> \
-o jsonpath='{range .status.conditions[?(@.type=="Paused")]}{.status} {.reason}{"\n"}{end}'
3. Remove the leftover jobs and pods
kubectl -n <namespace> delete job -l scaledjob.keda.sh/name=<scaledjob-name>

Read the queue before changing it. Confirm the queue name from the ScaledJob trigger rather than assuming it, because the name encodes the compute configuration the test used:

4. Confirm which queue the ScaledJob watches
kubectl -n <namespace> get scaledjob <scaledjob-name> \
-o jsonpath='{.spec.triggers[0].metadata.listName}{"\n"}'

Open an interactive Redis client. REDISCLI_AUTH keeps the password off the command line, where it would otherwise be visible in the pod spec and in process listings:

5. Open a Redis client
kubectl -n <namespace> exec -it deploy/dynamoai-api -- sh -c \
'REDISCLI_AUTH="$REDIS_PASSWORD" redis-cli -h "$REDIS_HOST" -p "$REDIS_PORT" --tls'

Inspect the queue, then clear only the orphaned entry:

6. Inspect, then remove the orphaned message
LLEN <queue-name>
LRANGE <queue-name> 0 -1
LREM <queue-name> 1 <entry>
Prefer LREM Over DEL

DEL <queue-name> discards every queued job on that queue, and any other test waiting on the same compute configuration must be resubmitted. Check LLEN first, and use LREM to remove the single orphaned entry whenever other work is pending.

7. Resume scaling
kubectl -n <namespace> annotate scaledjob <scaledjob-name> autoscaling.keda.sh/paused-

Worker pods should stop appearing once the queue reports 0. If they continue, the queue name is not the one you cleared; re-check it against the trigger in step 4.

Confirm Which Build a Worker Is Running​

Declared image tags can diverge from what is actually running, so confirm behaviour rather than configuration.

A worker on a build with Azure generation routing logs its decision at startup:

Check for the provider routing log line
kubectl -n <namespace> logs <worker-pod> | grep 'defaulting data generation'

If the line is absent on a deployment configured with DYNAMOEVAL_JUDGE_MODE=AZURE and AZURE_MODEL_NAME, the worker predates the routing change. Confirm the exact release with your DynamoAI deployment engineer.

Check every workload at once when a version bump may have applied unevenly:

List every DynamoAI image running in the namespace
kubectl -n <namespace> get deploy,statefulset,scaledjob -o json \
| grep -oE '"image": "[^"]*"' | sort -u

A partial upgrade, with the API on the new release and the workers still on the old one, produces behaviour that is difficult to attribute from the UI alone.