Skip to main content

Container Image Delivery

DynamoAI worker images are large, and the delivery path for them is the most common cause of a stalled first deployment or upgrade. Work through this page after the Environment Dependencies checks and before Deploy.

The short version: host DynamoAI images in a registry that stores them, size nodes for a multi-gigabyte image, and pre-pull for workers that scale from zero.

Image Characteristics​

Evaluation worker images are the largest the platform ships, and one layer dominates every transfer:

PropertyTypical value
Compressed sizeAbove 5 GiB
Largest single layerAbove 4 GiB
Extracted size on diskRoughly 1.5 to 2 times compressed
Layer reuse between releasesMinimal

Two consequences follow, and both are easy to plan for and expensive to discover late:

  1. Pull time is bounded by one layer. A single layer above 4 GiB cannot be split across connections, so additional bandwidth and parallelism do not shorten it.
  2. An upgrade transfers close to a whole image. Budget the full image size for every version change rather than a delta, including patch releases.

Registry​

The Container registry and artifact access requirements are defined in the dependencies registry. This section covers the delivery behaviour that requirement implies.

Registry Readiness0/4

Where a proxying remote repository cannot be avoided, raise its timeouts before validating. In JFrog Artifactory these live on the remote repository under Advanced:

SettingDefaultNeeded
Store Artifacts LocallyEnabledMust stay enabled, or nothing ever caches
Socket Timeout (ms)15000Raise to 300000; 15 seconds cannot move a 4 GiB layer
Missed Retrieval Cache Period (Sec)1800Caches negative lookups; set to 0 while validating
Unused Artifacts Cleanup Period (Hr)0A non-zero value evicts cached layers and the pull is paid again

A reverse proxy or load balancer in front of the registry has its own proxy_read_timeout and proxy_send_timeout, and those override the repository setting.

Node Capacity​

Node Capacity0/3

Version Upgrades​

Version Upgrades0/3
List every DynamoAI image running in the namespace
kubectl -n <namespace> get deploy,statefulset,scaledjob -o json \
| grep -oE '"image": "[^"]*"' | sort -u

Workers That Scale From Zero​

Evaluation workers run as KEDA ScaledJobs with minReplicaCount: 0. A node with no cached layers pays the full pull before the first job runs, and node replacement or autoscaling resets that cost.

Scale-From-Zero Readiness0/3

Pause the ScaledJob while warming the cache, and resume it afterwards:

Pause a ScaledJob during cache warming
kubectl -n <namespace> annotate scaledjob <scaledjob-name> \
autoscaling.keda.sh/paused="true" --overwrite

Network​

Network0/3

Verify Before Declaring the Environment Ready​

Pull the worker image onto a node with a plain pod and confirm it completes:

prepull-check.yaml
apiVersion: v1
kind: Pod
metadata:
name: prepull-check
namespace: <namespace>
spec:
restartPolicy: Never
containers:
- name: hold
image: <registry>/<project>/images/pen-testing:<tag>
command: ["sleep", "600"]
resources:
requests: { cpu: "100m", memory: "256Mi" }
Apply the check and read its events
kubectl apply -f prepull-check.yaml
kubectl -n <namespace> describe pod prepull-check | tail -20

A Successfully pulled image ... in Xm event is the pass condition. Use a plain pod rather than a Job or ScaledJob so that nothing deletes it mid-pull: a deletion cancels the transfer and reports context canceled, which looks identical to a timeout.

To confirm that bytes are moving, without needing a privileged pod:

Sample image filesystem usage on the node
kubectl get --raw "/api/v1/nodes/<node-name>/proxy/stats/summary" \
| python3 -c "import sys,json; print(json.load(sys.stdin)['node']['runtime']['imageFs']['usedBytes']/1024**3, 'GiB')"

Sample it twice, a minute apart. A flat value while a pull is in flight means the transfer is stalled, and the problem is the registry rather than the cluster.

Symptoms that appear after deployment, including context canceled on worker image pulls, are covered in Troubleshooting DynamoEval Workers.