Domain 3: Cloud Technology and Services

AWS AI/ML and Analytics Services (Task 3.7)

Amazon SageMaker Amazon Lex Amazon Kendra Amazon Polly Amazon Rekognition Amazon Textract Amazon Transcribe Amazon Translate Amazon Comprehend Amazon Athena Amazon Kinesis AWS Glue Amazon QuickSight Amazon EMR Amazon Redshift Amazon OpenSearch Amazon MSK
Exam Tip
Match service to task: chatbot = Lex; search documents = Kendra; text-to-speech = Polly; image/video analysis = Rekognition; extract from scanned docs = Textract; speech-to-text = Transcribe; translate = Translate; sentiment/entities = Comprehend; build/train your own model = SageMaker. Analytics: SQL on S3 files = Athena ($5/TB scanned); real-time streaming = Kinesis; ETL/catalog = Glue; dashboards = QuickSight; big data frameworks = EMR; data warehouse = Redshift; logs/search = OpenSearch.

AWS AI/ML and Analytics Services

Source: https://docs.aws.amazon.com/whitepapers/latest/aws-overview/analytics-and-ml-services.html

The AI/ML question is always "ready-made API or custom model?": the ready-made services are APIs you call with data; SageMaker is the platform where you BUILD models. Analytics questions are always "batch or streaming?" — then match the engine.

AI/ML Service Map — Match the Task

Task Service
Build/train/deploy your OWN models Amazon SageMaker
Chatbots / conversational interfaces Amazon Lex
Smart search across company documents Amazon Kendra
Text to speech (voice output) Amazon Polly
Speech to text (voice input) Amazon Transcribe
Image/video analysis (labels, faces, moderation) Amazon Rekognition
Extract text/tables/forms from scanned documents Amazon Textract
Language translation Amazon Translate
Sentiment, entities, key phrases in text Amazon Comprehend


How to remember: the direction of conversion names the service — Polly makes text SPEAK, Transcribe makes speech TEXT. Textract reads DOCUMENTS (OCR+), Rekognition sees IMAGES/VIDEO, Kendra SEARCHES your content, Lex CONVERSES, Comprehend READS sentiment, Translate converts LANGUAGE, SageMaker builds CUSTOM.

Real use-case: a customer-service suite: Lex chatbot front-end, Transcribe logs every phone call to text, Comprehend scores sentiment, Translate handles multilingual chats, and Kendra lets agents search the internal knowledge base naturally. None of this requires a data scientist — ready-to-call APIs. When you DO need custom models (predict churn from your own history), that is SageMaker (notebooks to prepare data, training jobs, endpoint deployment).

Gotchas & interview notes: the exam swaps input/output directions (Polly vs Transcribe) and image vs document (Rekognition vs Textract) — read the verb carefully. "Natural-language search over company documents" → Kendra, not OpenSearch (keyword search).

Analytics Services — Match the Data Problem

Problem Service Notes
Query data FILES in S3 with SQL, no servers Amazon Athena Serverless, $5/TB scanned
Ingest/stream real-time events Amazon Kinesis (Data Streams/Firehose/Analytics) Clickstreams, IoT, logs
ETL + data catalog AWS Glue Spark-based transforms; Glue Data Catalog indexes datasets
Petabyte data warehouse / BI SQL Amazon Redshift Columnar, massively parallel
Run Hadoop/Spark clusters Amazon EMR Managed big-data frameworks
Search + log analytics engine Amazon OpenSearch Service Full-text search, dashboards
Managed Kafka Amazon MSK Existing Kafka workloads
Dashboards / visual BI Amazon QuickSight Serverless BI, ML insights (anomaly detection)


The classic modern data stack (memorize the pipeline): IoT devices stream events → Kinesis Data Streams → Kinesis Data Firehose lands raw JSON in S3 → Glue crawls S3, builds the Data Catalog, and transforms to Parquet (cheaper to scan) → Athena for ad-hoc SQL → QuickSight dashboards for the business; heavy nightly jobs run on EMR.

How Kinesis works: Data Streams shards each ingest up to 1 MB/s and deliver 2 MB/s to consumers — consumers (Lambda, KCL apps) pull at their own pace; Firehose PUSHES to destinations (S3, Redshift, OpenSearch) with buffering/batching. Enhanced fan-out gives each consumer its own 2 MB/s pipe for heavy multi-consumer setups.

Real use-case: A fraud system must score transactions within 2 seconds of swipe: events ride Kinesis Data Streams; a consumer enriches and scores each event, pushing alerts to SNS. The same events, 30 minutes later, land in S3 via Firehose for nightly ML retraining — one pipeline serves real-time AND batch consumers.

Gotchas & interview notes: streaming vs batch is the first fork: seconds-level action → Kinesis; scheduled reports → Athena/Glue/EMR. Athena bills per TB SCANNED — partition by date and convert to Parquet to cut costs (a favorite cost scenario: "$5/TB becomes $0.50/TB with columnar formats"). Redshift = warehouse for SQL/BI on structured data; EMR = run your own Spark/Hadoop; OpenSearch = full-text search and log analytics — not a warehouse.

Streaming vs Batch — Quick Decision Rule

  • Data must be acted on within seconds (fraud alerts, live dashboards) → Kinesis (streaming)
  • Data processed on a schedule (nightly reports) → batch tooling: Athena/Glue/EMR

Cost Hook for Athena

Athena bills per TB scanned. Convert data to columnar formats (Parquet/ORC) and partition by date to cut scanned bytes — a favorite cost-optimization scenario.