At ~1,400 Workflows, EventSources at 512 Mi were OOMKilled; at ~2,000, with nodeStatusOffLoad false, the workflow controller at 2 GiB joined them. Label filters, CronWorkflow history limits, offload, and archive bounded the cache, not a message bus.
materialized='table' rewrites metadata.json with the full snapshot list. The coordinator parses that list into heap. File version in the 700s; JVM 2 GiB then 8 GiB. Expiry runs on Spark, not on Trino. Nessie GC is the wrong tool.
Nessie was medium-priority (600000); Trino coordinators were high-priority (1000000). The scheduler killed Nessie to place them. PDB is best-effort. Stock Nessie 0.103.3 could not set a PriorityClass. That is why the chart patch exists.
Iceberg needs a catalog; in-cluster MySQL had no TLS, no HA, and no backup. RocksDB is single-node, DocumentDB is not Mongo; we kept JDBC2, required TLS, and left the PVC until Aurora was trusted.
One coordinator for dashboards and dbt full-refresh is a bad idea. trino-main is humans; trino-etl is batch with task retry. dbt stays a container in an Argo workflow, not a Helm release.
Events exists so raw ingestion is not the same graph as semantic-layer dbt. Sensors on completion; producers stay a CronWorkflow, a label, a Sensor. The informer-cache OOM is a later story.
Argo Workflows was the data scheduler before Spark. The first job plane is Workflows + a standing Spark cluster: IRSA on the cluster, static keys on submit from `argo`, not the Spark Operator, not Argo CD, not Events yet.
First query path on an EKS cluster that already existed: Nessie + one Trino, JDBC2 on a PVC, CDK for the bucket, Helmfile for the two charts. It worked. No TLS, no HA, no Spark, no Argo jobs, no dbt.