Container Image Delivery
DynamoAI worker images are large, and the delivery path for them is the most common cause of a stalled first deployment or upgrade. Work through this page after the Environment Dependencies checks and before Deploy.
The short version: host DynamoAI images in a registry that stores them, size nodes for a multi-gigabyte image, and pre-pull for workers that scale from zero.
Image Characteristics
Evaluation worker images are the largest the platform ships, and one layer dominates every transfer:
| Property | Typical value |
|---|---|
| Compressed size | Above 5 GiB |
| Largest single layer | Above 4 GiB |
| Extracted size on disk | Roughly 1.5 to 2 times compressed |
| Layer reuse between releases | Minimal |
Two consequences follow, and both are easy to plan for and expensive to discover late:
- Pull time is bounded by one layer. A single layer above 4 GiB cannot be split across connections, so additional bandwidth and parallelism do not shorten it.
- An upgrade transfers close to a whole image. Budget the full image size for every version change rather than a delta, including patch releases.
Registry
The Container registry and artifact access requirements are defined in the dependencies registry. This section covers the delivery behaviour that requirement implies.
Where a proxying remote repository cannot be avoided, raise its timeouts before validating. In JFrog Artifactory these live on the remote repository under Advanced:
| Setting | Default | Needed |
|---|---|---|
| Store Artifacts Locally | Enabled | Must stay enabled, or nothing ever caches |
| Socket Timeout (ms) | 15000 | Raise to 300000; 15 seconds cannot move a 4 GiB layer |
| Missed Retrieval Cache Period (Sec) | 1800 | Caches negative lookups; set to 0 while validating |
| Unused Artifacts Cleanup Period (Hr) | 0 | A non-zero value evicts cached layers and the pull is paid again |
A reverse proxy or load balancer in front of the registry has its own proxy_read_timeout and proxy_send_timeout, and those override the repository setting.
Node Capacity
Version Upgrades
kubectl -n <namespace> get deploy,statefulset,scaledjob -o json \
| grep -oE '"image": "[^"]*"' | sort -u
Workers That Scale From Zero
Evaluation workers run as KEDA ScaledJobs with minReplicaCount: 0. A node with no cached layers pays the full pull before the first job runs, and node replacement or autoscaling resets that cost.
Pause the ScaledJob while warming the cache, and resume it afterwards:
kubectl -n <namespace> annotate scaledjob <scaledjob-name> \
autoscaling.keda.sh/paused="true" --overwrite
Network
Verify Before Declaring the Environment Ready
Pull the worker image onto a node with a plain pod and confirm it completes:
apiVersion: v1
kind: Pod
metadata:
name: prepull-check
namespace: <namespace>
spec:
restartPolicy: Never
containers:
- name: hold
image: <registry>/<project>/images/pen-testing:<tag>
command: ["sleep", "600"]
resources:
requests: { cpu: "100m", memory: "256Mi" }
kubectl apply -f prepull-check.yaml
kubectl -n <namespace> describe pod prepull-check | tail -20
A Successfully pulled image ... in Xm event is the pass condition. Use a plain pod rather than a Job or ScaledJob so that nothing deletes it mid-pull: a deletion cancels the transfer and reports context canceled, which looks identical to a timeout.
To confirm that bytes are moving, without needing a privileged pod:
kubectl get --raw "/api/v1/nodes/<node-name>/proxy/stats/summary" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['node']['runtime']['imageFs']['usedBytes']/1024**3, 'GiB')"
Sample it twice, a minute apart. A flat value while a pull is in flight means the transfer is stalled, and the problem is the registry rather than the cluster.
Symptoms that appear after deployment, including context canceled on worker image pulls, are covered in Troubleshooting DynamoEval Workers.