DynamoEval Workers
DynamoEval runs its attacks in KEDA ScaledJob workers that scale from zero. When a test never starts, or starts and fails immediately, the platform itself usually reports healthy: the HelmRelease is Ready, the API and UI are Running, and nothing in the namespace looks wrong. Each section below is keyed to the literal symptom, so match what you are seeing and skip the rest.
A worker takes one job, finishes it, and exits. The Job is then deleted after ttlSecondsAfterFinished, which defaults to 60 seconds, so the pod disappears before most people can read its logs. To keep a pod around while debugging, raise the value on the job you are investigating:
ttlSecondsAfterFinished: 3600
failedJobsHistoryLimit: 10
Lower it again afterwards, because retained pods hold node capacity.
Synthetic Data Generation Rejects the Configured Provider
A System Policy Compliance or Policy Jailbreak test fails within seconds:
guardrail compliant prompt generation failed for all generation models (<provider>/<model>).
Errors: <provider>/<model>: ValueError: Model <provider>/<model> not supported
This is not the provider rejecting the model, and the credentials are not wrong. The message is produced inside the worker before any network call. Synthetic data generation dispatches on the provider prefix of the model ID, and raises for any prefix it does not recognise. Every model under an unrecognised prefix produces the identical error.
| Prefix | Recognised for generation |
|---|---|
together_ai/ | Yes |
mistral/ | Yes |
openai/ | Yes |
gemini/ | Yes |
anthropic/ | Yes |
bedrock/ | Yes |
azure/, azure_ai/ | Yes, from platform release 3.26.7 |
Anything else, including vertex_ai/ | No |
Two attack types generate synthetic prompts before attacking, so only those two are affected:
| Attack | UI name | Affected |
|---|---|---|
guardrail_benchmark | System Policy Compliance | Yes |
policy_jailbreak | Policy Jailbreak | Yes |
static_jailbreak | Static Jailbreak | No |
adaptive_jailbreak | Adaptive Jailbreak | No |
Generation is a separate path from judging. Judges, attacker models, and the target system route through a component that supports a wider provider set, which is why the rest of a deployment works while only these two attacks fail.
Fix
Point generation at a provider the worker recognises and the cluster can reach. Leaving the generation models unset is not a fix on its own: the defaults are anthropic/claude-haiku-4-5, openai/gpt-4.1-mini, and gemini/gemini-2.5-flash, which an isolated network cannot reach.
Three environment variables set the generation models, one per fallback slot. The worker tries them in order until one succeeds:
DATA_GENERATION_ANTHROPIC_MODEL_IDDATA_GENERATION_OPENAI_MODEL_IDDATA_GENERATION_GEMINI_MODEL_ID
The variable names reflect the default provider in each slot, not a restriction: any recognised prefix is valid in any of the three.
- AWS
- Azure
- GCP
Set all three slots to a Bedrock model ID:
DATA_GENERATION_ANTHROPIC_MODEL_ID=bedrock/<model-id>
DATA_GENERATION_OPENAI_MODEL_ID=bedrock/<model-id>
DATA_GENERATION_GEMINI_MODEL_ID=bedrock/<model-id>
Bedrock generation reads four AWS variables, and all four are mandatory, including AWS_SESSION_TOKEN:
| Variable | Required |
|---|---|
AWS_ACCESS_KEY_ID | Yes |
AWS_SECRET_ACCESS_KEY | Yes |
AWS_REGION_NAME | Yes |
AWS_SESSION_TOKEN | Yes |
Generation reads these variables directly, so it does not pick up credentials from an instance role or from IAM Roles for Service Accounts. A missing variable fails the worker at configuration time rather than at the API call.
Setting DYNAMOEVAL_JUDGE_MODE to AWS_BEDROCK routes judges and attacker models to Bedrock, but it does not redirect generation. The three variables above are still required.
On platform release 3.26.7 and later, no additional variables are required. When DYNAMOEVAL_JUDGE_MODE is AZURE and AZURE_MODEL_NAME is set, the worker routes all three generation slots to that deployment automatically and logs the decision at startup:
Azure-only deployment: defaulting data generation models to the Azure deployment.
Set DATA_GENERATION_*_MODEL_ID to override.
Generation reuses the same credentials as the judge path, so a deployment already configured for Azure judges needs nothing further:
| Variable | Purpose |
|---|---|
AZURE_API_KEY | Comma-separated; entries after the first are treated as backups. On 3.26.8 and later, AZURE_ENTRA_AUTH can replace it; see the evaluation worker variables |
AZURE_API_BASE | Endpoint |
AZURE_API_VERSION | API version; ignored for azure_ai/ models |
Setting the three DATA_GENERATION_*_MODEL_ID variables explicitly is supported and harmless, but redundant.
Pointing DATA_GENERATION_*_MODEL_ID at an azure/ value on a build older than 3.26.7 converts a network failure into an immediate hard failure, because azure/ is not yet a recognised prefix. Upgrade the image first.
vertex_ai/ is not a recognised generation prefix, so generation cannot be routed to Vertex AI. Setting DYNAMOEVAL_JUDGE_MODE to GCP_VERTEX routes judges and attacker models to Vertex, but generation still resolves to the external provider defaults and fails in an isolated network.
Point the three generation slots at a recognised provider the cluster can reach:
DATA_GENERATION_ANTHROPIC_MODEL_ID=<prefix>/<model-id>
DATA_GENERATION_OPENAI_MODEL_ID=<prefix>/<model-id>
DATA_GENERATION_GEMINI_MODEL_ID=<prefix>/<model-id>
If no recognised provider is reachable from your environment, contact your DynamoAI deployment engineer before scheduling System Policy Compliance or Policy Jailbreak tests.
Two model-name formats coexist, and mixing them up produces errors that read as though the model itself is wrong:
| Variable | Format | Example |
|---|---|---|
AZURE_MODEL_NAME | Bare deployment name | <deployment> |
SYNDATA_MODEL_NAME, DATA_GENERATION_*_MODEL_ID | Provider-prefixed | azure/<deployment> |
A bare name where a prefix is expected fails with LLM Provider NOT provided.
Workers Keep Running the Previous Build After an Image Tag Change
The ScaledJob shows the new tag, the HelmRelease reports Ready, and worker behaviour is unchanged, including errors that the new build is supposed to fix.
Check the full image reference, not just the tag:
kubectl -n <namespace> get scaledjob <scaledjob-name> \
-o jsonpath='{.spec.jobTargetRef.template.spec.containers[0].image}{"\n"}'
If the output ends in @sha256:..., the values set both a tag and a digest. The chart concatenates them into repository:tag@digest, and container runtimes resolve by digest and ignore the tag entirely. Updating only the tag changes what the object displays and nothing about what runs.
Fix
Either remove the digest value so the tag is authoritative, or update the digest in the same change as the tag. Read the digest for the new tag from the registry:
curl -sI -u "<username>:<password>" \
-H "Accept: application/vnd.docker.distribution.manifest.v2+json" \
"https://<registry-host>/v2/<project>/images/pen-testing/manifests/<tag>" \
| grep -i docker-content-digest
pullPolicy Is Not the Key the Chart ReadsThe ScaledJob template reads image.imagePullPolicy. A value set as image.pullPolicy is silently ignored, so the pull policy falls back to the default. This has no effect while a digest is pinned, because digests are content-addressed, but it matters once the digest is removed.
Image Pull Fails With context canceled
Worker pods appear stuck for several minutes and then fail:
Failed to pull image "<registry>/<project>/images/pen-testing:<tag>":
rpc error: code = Canceled desc = failed to pull and unpack image
"<registry>/<project>/images/pen-testing:<tag>": context canceled
Error: ErrImagePull
Back-off pulling image "<registry>/<project>/images/pen-testing:<tag>"
Canceled is distinct from not found and unauthorized. The image exists and the credentials work; the transfer was aborted part-way. Kubelet cancels a pull that stops making measurable progress rather than waiting indefinitely.
Worker images are large, and one layer dominates the transfer. Two conditions commonly combine to stall it:
- The registry is a proxying remote and the tag is a cache miss. The proxy fetches the layer from upstream before serving it downstream, during which the node observes no progress.
- Several workers pull the same cold tag at once. KEDA creates a pod per queued job regardless of whether earlier pods have started, so bandwidth splits and every transfer slows toward the cancellation threshold.
Fix
Warm the registry cache with one deliberate pull from a machine that tolerates a long transfer, then let the workers pull from cache:
skopeo copy \
docker://<registry>/<project>/images/pen-testing:<tag> \
dir:/tmp/discard
A docker pull against the same reference works equally well. While warming, reduce maxReplicaCount on the affected job so the cluster is not attempting several cold pulls in parallel, then restore it.
Confirm the node can hold the extracted image if pulls still fail once the cache is warm:
kubectl describe node <node-name> | grep -A 6 Conditions
Sizing guidance and the pre-pull pattern that removes this failure mode are on the Container Image Delivery page.
Tests Stay at Preparing Resources With No Worker Pods
Preparing Resources means the attack is registered and every stage is still at 0%, so no worker has begun executing. Work down this chain, because each step has a different owner.
DynamoEval workers do not read the queue through KEDA's Redis scaler. They use the external scaler, so the path is:
API -> pushes job to Redis -> metrics-server reads queue length
-> reports over gRPC to KEDA -> KEDA scales the ScaledJob -> worker pod runs
A break anywhere in that chain looks identical from the UI.
kubectl -n <namespace> get pods | grep scaled-job
A pod in Pending phase is not necessarily unscheduled. A pod stays Pending until a container starts, which includes the whole image pull. Read its events before concluding anything:
kubectl -n <namespace> describe pod <pod-name> | tail -25
PodScheduled: True followed by Pulling means the pod has a node and is downloading. FailedScheduling names the constraint instead, commonly insufficient ephemeral-storage, because each worker requests a fixed reservation on top of a multi-gigabyte image.
kubectl -n <namespace> get scaledjobs
READY=True with ACTIVE=False and zero pods is the correct idle state, not a fault. ACTIVE=False while a test is genuinely queued means the scaler is reporting a queue length of zero; check that metrics-server is running and that it reads the same Redis database the API writes to. A mismatched database reports zero indefinitely.
READY=False means KEDA cannot reach the scaler at all. Verify the Service and its endpoints:
kubectl -n <namespace> get svc,endpoints dynamoai-metrics-server
kubectl -n keda get pods
kubectl -n keda logs deploy/keda-operator --tail=50 | grep -E 'external_scaler|Scaling Jobs'
Restarts on keda-operator, or repeated external_scaler errors reporting context canceled, indicate the operator is terminating mid-reconcile. It cannot then read queue lengths or write status conditions, so nothing scales. Identify the cause before restarting, because the restart clears the evidence:
kubectl -n keda describe pod <keda-operator-pod> | grep -A 15 "Last State"
kubectl -n keda logs <keda-operator-pod> --previous --tail=50
OOMKilled calls for a higher memory limit. A non-zero exit with lease-renewal messages in the previous logs points at leader election against a slow API server.
Worker Pods Are Recreated After a Test Is Cancelled
The test no longer appears in the UI, but KEDA keeps creating worker pods for its queue, often several, each replaced as it exits.
Cancelling a test removes its message from Redis only for attacks the API still finds in an IN_PROGRESS or QUEUED state, matched by job ID. If the attack record is already gone or in another state, cancellation returns NO_ACTION and the queued message is left behind. The scaler still sees a non-empty queue and keeps requesting workers, and no worker can complete the job because the test it belongs to no longer exists.
Fix
Stop the churn first, then clear the message.
kubectl -n <namespace> annotate scaledjob <scaledjob-name> \
autoscaling.keda.sh/paused="true" --overwrite
Confirm it took effect. The Paused condition reports ScaledJobPaused:
kubectl -n <namespace> get scaledjob <scaledjob-name> \
-o jsonpath='{range .status.conditions[?(@.type=="Paused")]}{.status} {.reason}{"\n"}{end}'
kubectl -n <namespace> delete job -l scaledjob.keda.sh/name=<scaledjob-name>
Read the queue before changing it. Confirm the queue name from the ScaledJob trigger rather than assuming it, because the name encodes the compute configuration the test used:
kubectl -n <namespace> get scaledjob <scaledjob-name> \
-o jsonpath='{.spec.triggers[0].metadata.listName}{"\n"}'
Open an interactive Redis client. REDISCLI_AUTH keeps the password off the command line, where it would otherwise be visible in the pod spec and in process listings:
kubectl -n <namespace> exec -it deploy/dynamoai-api -- sh -c \
'REDISCLI_AUTH="$REDIS_PASSWORD" redis-cli -h "$REDIS_HOST" -p "$REDIS_PORT" --tls'
Inspect the queue, then clear only the orphaned entry:
LLEN <queue-name>
LRANGE <queue-name> 0 -1
LREM <queue-name> 1 <entry>
LREM Over DELDEL <queue-name> discards every queued job on that queue, and any other test waiting on the same compute configuration must be resubmitted. Check LLEN first, and use LREM to remove the single orphaned entry whenever other work is pending.
kubectl -n <namespace> annotate scaledjob <scaledjob-name> autoscaling.keda.sh/paused-
Worker pods should stop appearing once the queue reports 0. If they continue, the queue name is not the one you cleared; re-check it against the trigger in step 4.
Confirm Which Build a Worker Is Running
Declared image tags can diverge from what is actually running, so confirm behaviour rather than configuration.
A worker on a build with Azure generation routing logs its decision at startup:
kubectl -n <namespace> logs <worker-pod> | grep 'defaulting data generation'
If the line is absent on a deployment configured with DYNAMOEVAL_JUDGE_MODE=AZURE and AZURE_MODEL_NAME, the worker predates the routing change. Confirm the exact release with your DynamoAI deployment engineer.
Check every workload at once when a version bump may have applied unevenly:
kubectl -n <namespace> get deploy,statefulset,scaledjob -o json \
| grep -oE '"image": "[^"]*"' | sort -u
A partial upgrade, with the API on the new release and the workers still on the old one, produces behaviour that is difficult to attribute from the UI alone.