AWS AI/ML and Analytics Services (Task 3.7)
AWS AI/ML and Analytics Services
Source: https://docs.aws.amazon.com/whitepapers/latest/aws-overview/analytics-and-ml-services.htmlThe AI/ML question is always "ready-made API or custom model?": the ready-made services are APIs you call with data; SageMaker is the platform where you BUILD models. Analytics questions are always "batch or streaming?" — then match the engine.
AI/ML Service Map — Match the Task
| Task | Service |
|---|---|
| Build/train/deploy your OWN models | Amazon SageMaker |
| Chatbots / conversational interfaces | Amazon Lex |
| Smart search across company documents | Amazon Kendra |
| Text to speech (voice output) | Amazon Polly |
| Speech to text (voice input) | Amazon Transcribe |
| Image/video analysis (labels, faces, moderation) | Amazon Rekognition |
| Extract text/tables/forms from scanned documents | Amazon Textract |
| Language translation | Amazon Translate |
| Sentiment, entities, key phrases in text | Amazon Comprehend |
How to remember: the direction of conversion names the service — Polly makes text SPEAK, Transcribe makes speech TEXT. Textract reads DOCUMENTS (OCR+), Rekognition sees IMAGES/VIDEO, Kendra SEARCHES your content, Lex CONVERSES, Comprehend READS sentiment, Translate converts LANGUAGE, SageMaker builds CUSTOM.
Real use-case: a customer-service suite: Lex chatbot front-end, Transcribe logs every phone call to text, Comprehend scores sentiment, Translate handles multilingual chats, and Kendra lets agents search the internal knowledge base naturally. None of this requires a data scientist — ready-to-call APIs. When you DO need custom models (predict churn from your own history), that is SageMaker (notebooks to prepare data, training jobs, endpoint deployment).
Gotchas & interview notes: the exam swaps input/output directions (Polly vs Transcribe) and image vs document (Rekognition vs Textract) — read the verb carefully. "Natural-language search over company documents" → Kendra, not OpenSearch (keyword search).
Analytics Services — Match the Data Problem
| Problem | Service | Notes |
|---|---|---|
| Query data FILES in S3 with SQL, no servers | Amazon Athena | Serverless, $5/TB scanned |
| Ingest/stream real-time events | Amazon Kinesis (Data Streams/Firehose/Analytics) | Clickstreams, IoT, logs |
| ETL + data catalog | AWS Glue | Spark-based transforms; Glue Data Catalog indexes datasets |
| Petabyte data warehouse / BI SQL | Amazon Redshift | Columnar, massively parallel |
| Run Hadoop/Spark clusters | Amazon EMR | Managed big-data frameworks |
| Search + log analytics engine | Amazon OpenSearch Service | Full-text search, dashboards |
| Managed Kafka | Amazon MSK | Existing Kafka workloads |
| Dashboards / visual BI | Amazon QuickSight | Serverless BI, ML insights (anomaly detection) |
The classic modern data stack (memorize the pipeline): IoT devices stream events → Kinesis Data Streams → Kinesis Data Firehose lands raw JSON in S3 → Glue crawls S3, builds the Data Catalog, and transforms to Parquet (cheaper to scan) → Athena for ad-hoc SQL → QuickSight dashboards for the business; heavy nightly jobs run on EMR.
How Kinesis works: Data Streams shards each ingest up to 1 MB/s and deliver 2 MB/s to consumers — consumers (Lambda, KCL apps) pull at their own pace; Firehose PUSHES to destinations (S3, Redshift, OpenSearch) with buffering/batching. Enhanced fan-out gives each consumer its own 2 MB/s pipe for heavy multi-consumer setups.
Real use-case: A fraud system must score transactions within 2 seconds of swipe: events ride Kinesis Data Streams; a consumer enriches and scores each event, pushing alerts to SNS. The same events, 30 minutes later, land in S3 via Firehose for nightly ML retraining — one pipeline serves real-time AND batch consumers.
Gotchas & interview notes: streaming vs batch is the first fork: seconds-level action → Kinesis; scheduled reports → Athena/Glue/EMR. Athena bills per TB SCANNED — partition by date and convert to Parquet to cut costs (a favorite cost scenario: "$5/TB becomes $0.50/TB with columnar formats"). Redshift = warehouse for SQL/BI on structured data; EMR = run your own Spark/Hadoop; OpenSearch = full-text search and log analytics — not a warehouse.
Streaming vs Batch — Quick Decision Rule
- Data must be acted on within seconds (fraud alerts, live dashboards) → Kinesis (streaming)
- Data processed on a schedule (nightly reports) → batch tooling: Athena/Glue/EMR
Cost Hook for Athena
Athena bills per TB scanned. Convert data to columnar formats (Parquet/ORC) and partition by date to cut scanned bytes — a favorite cost-optimization scenario.