← Selected work
AWS Serverless · Event-Driven · Project 8

Document Intelligence Pipeline

Upload any PDF — the pipeline automatically extracts text and uses Bedrock Claude to analyze it. Built entirely on AWS serverless services: S3, Lambda, Bedrock, DynamoDB, and API Gateway. No server running 24/7. Everything triggered by events. Infrastructure defined as code with AWS CDK.

AWS Lambda · S3 · Bedrock · DynamoDB · API Gateway · CDK · Node.js · React GitHub ↗
New Concepts This Project
AWS Lambda — serverless function, runs only when triggered, no server needed.
S3 Event Trigger — upload PDF → S3 automatically fires Lambda. No manual trigger.
AWS Bedrock — same Claude AI but running inside your AWS account. Data never leaves AWS.
DynamoDB — NoSQL database, perfect for Lambda (no persistent connections needed).
API Gateway — REST endpoints without a server, replaces Express app.listen().
AWS CDK — infrastructure as code, deploy everything with one command.
Event-driven architecture — services react to events, nothing runs 24/7.

What It Does

You upload invoice.pdf
        ↓
S3 stores the file
        ↓
S3 fires Lambda automatically (event trigger)
        ↓
Lambda extracts text using pdf-parse
        ↓
Bedrock Claude analyzes the text
        ↓
Results saved to DynamoDB:
  "This is an invoice from Acme Corp
   Amount: $4,500
   Due date: Aug 30, 2026
   Action: Payment required"
        ↓
React UI shows analysis ✅

No button to click to "process"
No server waiting for requests
Upload happens → everything starts automatically ✅

Why AWS? — What's Different From Projects 1-7

Projects 1-7 (Express server):
  Node.js server running 24/7
  React → http://localhost:3006 → Express → Anthropic API

  Problems at scale:
    Server always running → paying even when idle ❌
    One server → can't handle 1000 users at once ❌
    You manage everything → patches, restarts, scaling ❌

Project 8 (AWS serverless):
  No server running 24/7 ✅
  Lambda runs ONLY when PDF uploaded ✅
  AWS handles scaling automatically ✅
  Pay only when code actually runs ✅
  1 user or 1 million → same setup ✅

Complete Architecture

React UI (browser)
  │
  ├── Upload PDF
  │     ↓
  │   POST /upload-url → API Gateway → Lambda
  │     ↓ presigned S3 URL returned
  │   PUT directly to S3 (browser → S3)
  │     ↓ S3 event trigger (automatic!)
  │   AWS Lambda — processor
  │     ↓ reads PDF from S3
  │   pdf-parse — extract text
  │     ↓ raw text
  │   AWS Bedrock Claude — analyze
  │     ↓ structured JSON analysis
  │   AWS DynamoDB — save results
  │
  └── View Results
        ↓
      GET /document → API Gateway → Lambda
        ↓
      DynamoDB — fetch results
        ↓
      React UI — display analysis ✅

AWS Services — What Each One Does

📦 Amazon S3 — Simple Storage Service
What it is:
  File storage on AWS — like Google Drive for code/apps
  Stores files permanently, globally available

What it does in P8:
  Stores uploaded PDF files permanently
  Triggers Lambda automatically when file arrives

Why not store locally?
  Local file → server restarts → file gone ❌
  S3 → permanent storage → never lost ✅
  S3 → triggers Lambda automatically ✅

Real world analogy:
  S3 = a smart mailbox
  When mail arrives → mailbox rings a bell (event)
  Bell wakes up Lambda ✅

Setup command:
  aws s3 mb s3://your-bucket-name --region us-east-1

CORS required for browser uploads:
  aws s3api put-bucket-cors \
    --bucket your-bucket-name \
    --cors-configuration file://s3-cors.json

Cost: First 5GB free, then $0.023/GB
⚡ AWS Lambda — Serverless Functions
What it is:
  A function that runs on AWS
  No server needed — AWS manages everything

Compare with Express (what you know):

  Express server (P1-P7):
    const app = express();
    app.post("/process", handler);
    app.listen(3006);  // always running ❌

  Lambda (P8):
    export const handler = async (event) => {
      // same logic here
      return { statusCode: 200 };
    };
    // No app.listen()
    // AWS calls handler() for you ✅

Key difference:
  Express → runs 24/7 waiting for requests
  Lambda  → sleeps → wakes when S3 uploads PDF
            runs → sleeps again ✅

How Lambda executes without a server:
  You deploy code (zip file) to AWS
  AWS stores it — NOT running
        ↓
  Event happens (PDF uploaded to S3)
        ↓
  AWS spins up container:
    → Unzips your code
    → Starts Node.js 20 runtime
    → Loads index.mjs
    → Calls handler(event)
        ↓
  Your code runs
        ↓
  Container sleeps (stays warm 15 min)

Cold start vs Warm start:
  Cold: first request → ~1-2 seconds (container spin up)
  Warm: reuses container → ~100ms ✅

Cost: First 1 million requests FREE every month
🧠 AWS Bedrock — Claude AI Inside AWS
What it is:
  Same Claude AI we use in P1-P7
  But running INSIDE your AWS account

Compare:
  Projects 1-7 — Anthropic public API:
    Your code → api.anthropic.com (public internet)
    Anthropic servers can see your data ⚠️

  Project 8 — AWS Bedrock:
    Your code → AWS Bedrock (inside AWS)
    Data never leaves your AWS account ✅
    HIPAA compliant ✅
    Enterprise ready ✅

Why this matters for enterprise:
  Companies with sensitive data (medical, legal, financial)
  CANNOT send data to public Anthropic API
  Must use Bedrock or Azure OpenAI ✅

Model ID:
  us.anthropic.claude-sonnet-4-6
  ↑ "us." prefix = US inference profile
  Required for on-demand usage (not just base model ID)

Common error:
  "Model not supported for on-demand throughput"
  Fix: add "us." prefix to model ID ✅
🗄️ AWS DynamoDB — NoSQL Database
What it is:
  AWS managed NoSQL database
  Key-value store — like a giant JSON object
  No fixed schema — store any shape of data

Compare with Postgres (what you know):

  Postgres (P3, P5, P7):
    Fixed columns, SQL queries
    Needs persistent TCP connection
    Connection overhead with Lambda ❌

  DynamoDB (P8):
    Flexible JSON documents
    HTTP-based — no persistent connection
    Perfect for Lambda ✅

Why DynamoDB for Lambda (not Postgres)?
  Lambda starts/stops constantly
  Opening new Postgres connection every time = slow ❌
  DynamoDB = HTTP call = no connection overhead ✅

What we store:
  {
    documentId:    "uuid-123",
    fileName:      "invoice.pdf",
    status:        "completed",
    extractedText: "Invoice from Acme...",
    analysis: {
      documentType: "invoice",
      summary:      "Invoice for $4,500...",
      amounts:      ["$4,500"],
      dueDate:      "Aug 30, 2026"
    }
  }

Setup command:
  aws dynamodb create-table \
    --table-name document-results \
    --attribute-definitions AttributeName=documentId,AttributeType=S \
    --key-schema AttributeName=documentId,KeyType=HASH \
    --billing-mode PAY_PER_REQUEST

Cost: First 25GB free forever
🌐 API Gateway — REST Endpoints Without Server
What it is:
  Creates HTTP endpoints that trigger Lambda
  Replaces Express routes completely

Compare:
  Express (P1-P7):
    app.get("/document/:id", handler);
    → http://localhost:3006/document/123

  API Gateway (P8):
    GET /document/{id} → triggers Lambda
    → https://xxx.execute-api.us-east-1.amazonaws.com/prod/document/123

React calls AWS URL instead of localhost
Same logic inside Lambda ✅

Routes we created:
  POST /upload-url    → get presigned S3 URL
  GET  /document      → list all documents
  GET  /document/{id} → get one document

Why presigned URL for upload?
  React cannot upload directly to S3 (CORS + auth)

  Solution:
    1. React asks API: "give me upload permission"
    2. API Lambda generates secure S3 URL (5 min expiry)
    3. React uploads directly to S3 using that URL
    4. S3 triggers processor Lambda automatically ✅

IAM Role — Why Lambda Needs Permission

Lambda needs permission to use other AWS services.
Without permission → "Access Denied" error ❌

We created: document-intelligence-role

Permissions attached:
  AWSLambdaBasicExecutionRole  → write logs to CloudWatch
  AmazonS3ReadOnlyAccess       → read PDF from S3
  AmazonBedrockFullAccess      → call Claude via Bedrock
  AmazonDynamoDBFullAccess     → read/write results

Real world analogy:
  Lambda = new employee
  IAM Role = employee badge
  Badge gives access to specific rooms:
    S3 room (read files) ✅
    Bedrock room (call AI) ✅
    DynamoDB room (save results) ✅

Production approach — least privilege:
  Only grant EXACT permissions needed
  NOT AdministratorAccess ❌
  Specific actions on specific resources ✅

Setup commands:
  aws iam create-role \
    --role-name document-intelligence-role \
    --assume-role-policy-document file://trust-policy.json

  aws iam attach-role-policy \
    --role-name document-intelligence-role \
    --policy-arn arn:aws:iam::aws:policy/AmazonBedrockFullAccess

S3 Event Trigger — How It Works

Configured S3 to watch for PDF uploads:
  "When any .pdf file is uploaded
   → automatically call Lambda"

Only .pdf files trigger Lambda
Other files (.jpg, .docx) → ignored ✅

The trigger flow:
  You upload invoice.pdf to S3
        ↓
  S3 detects: new .pdf file!
        ↓
  S3 sends event to Lambda:
    {
      Records: [{
        s3: {
          bucket: { name: "mahesh-document-intelligence" },
          object: { key: "uploads/uuid-invoice.pdf" }
        }
      }]
    }
        ↓
  Lambda wakes up → processes PDF ✅

Setup commands:
  # Give S3 permission to invoke Lambda:
  aws lambda add-permission \
    --function-name document-intelligence \
    --statement-id s3-trigger \
    --action lambda:InvokeFunction \
    --principal s3.amazonaws.com \
    --source-arn arn:aws:s3:::your-bucket

  # Connect S3 to Lambda:
  aws s3api put-bucket-notification-configuration \
    --bucket your-bucket \
    --notification-configuration file://s3-notification.json

Lambda Code — index.mjs

// No Express needed — just a plain async function

export const handler = async (event) => {

  // 1. Get PDF details from S3 event
  const bucket   = event.Records[0].s3.bucket.name;
  const s3Key    = event.Records[0].s3.object.key;
  const fileName = s3Key.split("/").pop();

  // 2. Save "processing" status to DynamoDB
  await saveInitialRecord(documentId, fileName);

  // 3. Download PDF from S3
  const s3Response = await s3.send(new GetObjectCommand({ Bucket: bucket, Key: s3Key }));
  const chunks = [];
  for await (const chunk of s3Response.Body) chunks.push(chunk);
  const buffer = Buffer.concat(chunks);

  // 4. Extract text using pdf-parse (free — no Textract needed)
  const pdfData = await pdf(buffer);
  const text    = pdfData.text;

  // 5. Analyze with Bedrock Claude
  const response = await bedrock.send(new InvokeModelCommand({
    modelId: "us.anthropic.claude-sonnet-4-6",
    body: Buffer.from(JSON.stringify({
      anthropic_version: "bedrock-2023-05-31",
      max_tokens: 1000,
      messages: [{ role: "user", content: prompt }]
    })),
    contentType: "application/json",
    accept:      "application/json",
  }));

  // 6. Save results to DynamoDB
  await saveResults(documentId, text, analysis);

  return { statusCode: 200 };
  // No res.json() — just return object ✅
  // No app.listen() — AWS calls this for you ✅
};

AWS CDK — Infrastructure as Code

Why CDK instead of manual CLI commands?

Manual approach (what we started with):
  aws s3 mb s3://bucket-name
  aws dynamodb create-table ...
  aws lambda create-function ...
  aws apigateway create-rest-api ...
  → Run 20+ commands every time ❌
  → Easy to make mistakes ❌
  → Not version controlled ❌

CDK approach (production standard):
  cdk deploy  ← one command creates everything ✅
  Version controlled in GitHub ✅
  Reproducible across environments ✅
  Readable TypeScript code ✅

CDK stack defines everything:
  const bucket = s3.Bucket.fromBucketName(...);  // import existing
  const table  = dynamodb.Table.fromTableName(...); // import existing

  const processor = new lambda.Function(this, "Processor", {
    runtime:  lambda.Runtime.NODEJS_20_X,
    handler:  "index.handler",
    code:     lambda.Code.fromAsset("../lambda"),
    timeout:  cdk.Duration.seconds(120),
    environment: { DYNAMODB_TABLE: table.tableName },
  });

  const api = new apigateway.RestApi(this, "Api", {
    defaultCorsPreflightOptions: {
      allowOrigins: apigateway.Cors.ALL_ORIGINS,
    },
  });

Deploy:
  cdk bootstrap aws://ACCOUNT/us-east-1  // once per account
  npm run build
  cdk diff    // preview changes
  cdk deploy  // deploy ✅

Output:
  DocumentIntelligenceStack.ApiUrl = https://xxx.amazonaws.com/prod/

Errors We Hit and How We Fixed Them

Error 1 — Textract subscription required:
  AWS Textract costs money for new accounts
  Fix: use pdf-parse npm package instead (free)
  import pdf from "pdf-parse/lib/pdf-parse.js";

Error 2 — Model not supported for on-demand:
  Wrong: anthropic.claude-sonnet-4-6
  Right: us.anthropic.claude-sonnet-4-6
  Fix: add "us." prefix = US inference profile

Error 3 — CORS error on PDF upload:
  Browser blocks S3 upload (no CORS config)
  Fix: add CORS rules to S3 bucket
  aws s3api put-bucket-cors --bucket ... --cors-configuration file://cors.json

Error 4 — S3 bucket already exists in CDK:
  Created bucket manually first, CDK tries to create again
  Fix: import existing instead of creating new
  const bucket = s3.Bucket.fromBucketName(this, "Bucket", "existing-name");

Error 5 — Overlapping S3 notifications:
  Manual trigger + CDK trigger = conflict
  Fix: remove addEventNotification from CDK stack
  Keep the manually created trigger

Error 6 — JSON malformed in PowerShell:
  PowerShell escapes JSON differently
  Fix: save JSON to file first, use file://
  Set-Content -Path config.json -Value '{...}'
  aws ... --configuration file://config.json

Error 7 — 403 on presigned URL upload:
  Lambda role doesn't have S3 PutObject permission
  Fix: attach S3 write policy to CDK Lambda role
  aws iam attach-role-policy --role-name CDK-ROLE --policy-arn ...

Security — API URL in React

Problem — hardcoded API URL in React:
  const API_URL = "https://xxx.execute-api.amazonaws.com/prod";
  → Visible in browser DevTools
  → Anyone can call your API ❌

Fix — environment variables:
  // App.jsx:
  const API_URL = import.meta.env.VITE_API_URL;

  // client/.env (never commit to GitHub):
  VITE_API_URL=https://xxx.execute-api.amazonaws.com/prod

  // client/.env.example (safe to commit):
  VITE_API_URL=your-api-gateway-url

Is API Gateway URL a secret?
  URL alone: NOT a secret ✅
  Without auth: anyone can call it ⚠️
  With auth: URL useless without token ✅

Production auth options:
  API Keys:  simple, add x-api-key header
  Cognito:   full user auth with JWT tokens
  IAM auth:  AWS-level authentication

For portfolio/demo:
  URL in .env is fine ✅
  Add note: "Production would add Cognito auth"

What NEVER goes in code or GitHub:
  ❌ AWS Access Key ID
  ❌ AWS Secret Access Key
  ❌ API keys or tokens
  ❌ .env files

Lambda vs Express — Full Comparison

Feature           Express (P1-P7)        Lambda (P8)
────────────────────────────────────────────────────────
Always running    YES ❌                 NO ✅
Server needed     YES ❌                 NO ✅
Triggered by      HTTP request only      S3, HTTP, timer, anything
Scales            Manual ❌              Auto ✅
Cost              Pay 24/7               Pay per execution
Code entry        app.listen(3006)       export const handler
Input             req.body               event object
Output            res.json()             return {}
Localhost test    YES ✅                 NO (deploy to test)
Port number       3006, 3007 etc         None
Deployment        node index.js          zip → aws lambda update

DynamoDB vs Postgres — When to Use Each

Use Postgres when:
  → Complex SQL queries needed ✅
  → Relationships between tables ✅
  → pgvector for embeddings ✅
  → Always-on server ✅
  → Used in: P3, P5, P7

Use DynamoDB when:
  → Serverless (Lambda) ✅
  → Simple key-value lookups ✅
  → Need massive auto-scale ✅
  → No persistent connection ✅
  → Used in: P8

Why DynamoDB for Lambda:
  Lambda opens/closes constantly
  Postgres: persistent TCP connection = overhead per invocation ❌
  DynamoDB: HTTP call = no connection management ✅

Bedrock vs Anthropic API

Anthropic API (P1-P7):
  Your code → public internet → Anthropic servers
  Anthropic can see your data ⚠️
  Fine for: demos, personal projects, non-sensitive data

AWS Bedrock (P8):
  Your code → private AWS network → Bedrock
  Data stays in YOUR AWS account ✅
  HIPAA compliant ✅
  SOC2 compliant ✅
  Enterprise ready ✅

Same Claude model, same responses
Different data residency and compliance

When companies MUST use Bedrock:
  Medical records → HIPAA requirement
  Legal documents → client confidentiality
  Financial data  → regulatory compliance
  Any PII        → GDPR considerations

Deploying Lambda — No Localhost

Projects 1-7 development:
  Edit code → node index.js → test at localhost ✅
  Instant feedback loop ✅

Lambda development:
  Edit code → zip → upload to AWS → test
  ~30 seconds per change ⚠️

No localhost because:
  Lambda needs real S3, Bedrock, DynamoDB
  These only exist on AWS
  Cannot simulate locally (without LocalStack)

Deploy commands:
  # Zip:
  Compress-Archive -Path * -DestinationPath ../lambda.zip -Force

  # First deploy:
  aws lambda create-function \
    --function-name document-intelligence \
    --zip-file fileb://lambda.zip ...

  # Update existing:
  aws lambda update-function-code \
    --function-name document-intelligence \
    --zip-file fileb://lambda.zip

Production solution — CI/CD:
  git push → GitHub Actions → auto zip → auto deploy
  Same 30 seconds but fully automated ✅

CloudWatch — Seeing Lambda Logs

Projects 1-7:
  console.log("Processing...") → terminal window ✅

Lambda:
  console.log("Processing...") → CloudWatch logs

How to view:
  AWS Console → CloudWatch
  → Log groups
  → /aws/lambda/document-intelligence
  → Click latest log stream
  → See all logs ✅

Our logs show:
  📥 Event received: {bucket, key, size}
  📄 Processing: invoice.pdf from mahesh-document-intelligence
  💾 Saving initial record to DynamoDB...
  🔍 Extracting text from PDF...
  ✅ Extracted 1,234 characters
  🧠 Analyzing with Bedrock Claude...
  ✅ Analysis complete
  💾 Saving results to DynamoDB...
  ✅ Document processed successfully

Example Claude Analysis Output

Input: Upload invoice.pdf

Claude's analysis saved to DynamoDB:
{
  "documentType": "invoice",
  "summary": "Invoice from Acme Corp for software
              development services totaling $4,500,
              due August 30, 2026.",
  "keyFields": {
    "vendor":      "Acme Corp",
    "invoiceNo":   "INV-2026-0892",
    "amount":      "$4,500",
    "dueDate":     "August 30, 2026",
    "services":    "Software development"
  },
  "importantDates": ["August 30, 2026"],
  "amounts":        ["$4,500"],
  "actionItems":    ["Payment due August 30, 2026"]
}

Cost Breakdown

Service        Free Tier              Our Usage
─────────────────────────────────────────────────────
S3             5GB/month              < 1GB ✅
Lambda         1M requests/month      < 1000 ✅
Bedrock        Pay per token          ~$0.003/doc
DynamoDB       25GB forever           < 1GB ✅
API Gateway    1M calls/month         < 1000 ✅
CloudWatch     Basic logs free        ✅
─────────────────────────────────────────────────────
Total for demo: ~$0-2/month ✅

Production Enhancements

This project is demo grade.
Production would add:

Security:
  → Cognito auth on API Gateway (user login)
  → WAF (Web Application Firewall)
  → Least privilege IAM policies
  → API keys or JWT for all endpoints
  → Environment variables (not hardcoded URLs)

Reliability:
  → Dead Letter Queue (SQS) for failed Lambdas
  → CloudWatch alarms on error spikes
  → AWS X-Ray for request tracing
  → Multi-AZ DynamoDB

CI/CD:
  → GitHub Actions → cdk deploy on push
  → Separate dev / staging / prod accounts
  → Automated tests before deploy
  → Rollback on failure

Local Development:
  → LocalStack — full AWS simulation locally
  → aws-sdk-client-mock for unit tests
  → Docker Compose for local services
  → No more "deploy to test" cycle

Cost Optimization:
  → S3 lifecycle rules (delete after 90 days)
  → DynamoDB TTL (auto-expire old records)
  → Lambda reserved concurrency limits
  → CloudWatch log retention policies

Docker — Containerizing Lambda

Why Docker for Lambda?

Without Docker (zip deployment):
  zip lambda/ → upload to AWS
  Different environments = different behavior ❌
  Hard to reproduce bugs ❌

With Docker (container):
  Same container runs locally + on AWS ✅
  Consistent environment always ✅
  Version tagged images = easy rollback ✅
  Up to 10GB image (vs 250MB zip) ✅

Our Dockerfile:
  FROM public.ecr.aws/lambda/nodejs:20
  ← AWS official Lambda image, pre-configured

  COPY package.json package-lock.json ./
  RUN npm ci --omit=dev
  ← Production only, no dev dependencies
  ← Smaller image = faster cold start ✅

  COPY index.mjs ./
  COPY api.mjs ./

  CMD ["index.handler"]
  ← Default handler, override per Lambda function

Build:
  docker build -t document-intelligence .

Why FROM public.ecr.aws/lambda/nodejs:20?
  → AWS official image
  → Already configured for Lambda runtime
  → No manual Lambda bootstrap needed ✅

ECR — AWS Container Registry

What is ECR?
  Elastic Container Registry
  = AWS private Docker Hub
  = Where Lambda pulls your Docker image from

Why ECR not Docker Hub?
  Docker Hub: public internet ❌
  ECR:        inside AWS account ✅
              faster pulls for Lambda ✅
              integrated with IAM auth ✅
              private by default ✅

Setup:
  aws ecr create-repository \
    --repository-name document-intelligence \
    --region us-east-1

  Returns: repositoryUri
  571600846443.dkr.ecr.us-east-1.amazonaws.com/document-intelligence

Three steps to push image:
  # Step 1: Login
  aws ecr get-login-password --region us-east-1 \
    | docker login --username AWS \
      --password-stdin 571600846443.dkr.ecr.us-east-1.amazonaws.com

  # Step 2: Tag
  docker tag document-intelligence:latest \
    571600846443.dkr.ecr.us-east-1.amazonaws.com/document-intelligence:latest

  # Step 3: Push
  docker push \
    571600846443.dkr.ecr.us-east-1.amazonaws.com/document-intelligence:latest

Versioned images in ECR:
  :latest   → always newest
  :abc123   → specific git commit SHA
  → Rollback to any version instantly ✅

GitHub Actions — CI/CD Pipeline

What is CI/CD?
  CI = Continuous Integration
       Every push → build + test automatically
  CD = Continuous Deployment
       Every merge to main → auto deploy ✅

Before CI/CD (manual every time):
  Edit code
  → zip lambda/
  → aws lambda update-function-code ← manual ❌
  → Easy to forget steps ❌
  → No consistency ❌

After CI/CD (automated):
  Edit code → git push
  → Everything happens automatically ✅
  → Same steps every single time ✅

Our workflow (.github/workflows/deploy.yml):
Trigger: push to main branch

  Step 1:  Checkout code
  Step 2:  Setup Node.js 20
  Step 3:  npm ci (install deps)
  Step 4:  Configure AWS credentials
           → reads from GitHub Secrets (safe!)
  Step 5:  Login to ECR
  Step 6:  Build Docker image
           docker build -t ECR_URI:GIT_SHA .
  Step 7:  Push to ECR
           docker push ECR_URI:GIT_SHA
           docker push ECR_URI:latest
  Step 8:  Zip Lambda code
  Step 9:  Deploy processor Lambda
           aws lambda update-function-code
  Step 10: Wait for Lambda update complete
  Step 11: Deploy API Lambda
  Step 12: Verify both Lambdas Active ✅

Total time: ~2-3 minutes
Zero manual steps after git push ✅

GitHub Secrets — Safe Credentials in CI/CD

Problem:
  CI/CD needs AWS credentials to deploy
  Repo is public — can't hardcode credentials ❌

Solution — GitHub Secrets:
  Encrypted at rest ✅
  Never shown in logs (shown as ***) ✅
  Only accessible during workflow run ✅
  Even repo owner can't read values back ✅

Setup:
  github.com/repo → Settings
  → Secrets and variables → Actions
  → New repository secret

  Name:  AWS_ACCESS_KEY_ID
  Value: AKIA... (your IAM key)

  Name:  AWS_SECRET_ACCESS_KEY
  Value: xxxx... (your IAM secret)

Used in workflow:
  - uses: aws-actions/configure-aws-credentials@v4
    with:
      aws-access-key-id: ${{ secrets.AWS_ACCESS_KEY_ID }}
      aws-secret-access-key: ${{ secrets.AWS_SECRET_ACCESS_KEY }}

In GitHub Actions logs:
  Configuring credentials...
  aws-access-key-id: ***  ← always masked ✅

Safe to put in public repo:
  ✅ Workflow file with secret names
  ✅ Region and account ID
  ❌ Never actual credential values
  ❌ Never .env files

Full Automated Flow — Code to Production

You fix a bug in index.mjs
        ↓
git commit -m "fix: handle empty PDF gracefully"
git push origin main
        ↓
GitHub detects push to main branch
        ↓
GitHub Actions workflow triggers automatically
        ↓
Ubuntu runner spins up:
  ✅ Checkout your code
  ✅ npm install (production deps)
  ✅ Configure AWS (from encrypted secrets)
  ✅ Login to ECR
  ✅ docker build → new versioned image
  ✅ docker push → stored in ECR
     :abc123def (commit SHA)
     :latest
  ✅ zip lambda/ → lambda.zip
  ✅ aws lambda update-function-code
     document-intelligence ← processor
  ✅ aws lambda wait function-updated
  ✅ aws lambda update-function-code
     document-intelligence-api ← API
  ✅ aws lambda wait function-updated
  ✅ Verify both Lambdas Active
        ↓
~2-3 minutes from push to production ✅
Zero manual steps ✅
Consistent every time ✅
Full audit trail in GitHub Actions ✅
Why This Project Matters for Your Career

Every enterprise company uses AWS. Lambda, S3, DynamoDB, and API Gateway are the most asked-about services in interviews. Before P8 you could only talk about Cloudflare and Anthropic. After P8 you can speak to AWS serverless architecture from hands-on experience — not theory. You hit real errors (Textract costs, model ID prefix, CORS, IAM permissions) and fixed them. That real-world debugging experience is what interviewers want to hear about. The event-driven pattern (S3 → Lambda → DynamoDB) is used by Netflix, Airbnb, and thousands of startups to process millions of files daily.