AI for the cloud engineer: the infrastructure perspective
How AI workloads behave differently from traditional cloud applications, and what that means for the engineers who run the infrastructure.
- How AI workloads differ from traditional applications: compute profiles, latency, cost structures
- GPU vs CPU: selecting the right instance type for AI workloads on AWS and Azure
- The AI cloud stack: managed services, build vs buy, and where the abstraction layer sits
- Key vocabulary for cloud engineers: tokens, inference, embeddings, RAG, fine-tuning
Deploy an LLM inference endpoint on AWS Bedrock and benchmark latency and cost.
An infrastructure assessment of how your current cloud environment would need to change to support an AI workload.

