Christian Santiago

AWS Certified Developer

Software engineer building generative AI infrastructure on AWS: agents, RAG, inference serving, evaluation, and guardrails.

I came up through infrastructure monitoring, incident response, and the data engineering around it, working directly with the clients and stakeholders who depended on the output. Lately I have been building AWS infrastructure in Go, Python, TypeScript, and Terraform for generative AI workloads, and I am focused on agentic RAG and document extraction.

About

A bit about me.

A short professional summary and some personal context.

Professional

My work has spanned both ends of the scale, and each end teaches something the other cannot. Enterprise infrastructure, where a change to a shared system has a blast radius and you learn to respect it. Small and mid-sized businesses, where there is no layer between you and the client who needs the answer, so you gather the requirement, build the pipeline, and present the result yourself.

I spent four years in an enterprise environment on infrastructure monitoring and incident response, and the data engineering around it: thresholds and event-driven alert routing on the detection side, telemetry pipelines and dimensional models on the side that turned resolved incidents into failure patterns. All of it through an on-premise to hybrid AWS migration. Since then, systems integration, ETL pipelines, BI reporting for leadership, and freelance application development.

The through line in my work is system design, which has less to do with knowing services than with asking what tradeoffs a decision makes. Reliability against cost. Simplicity against flexibility. Shipping now against unwinding it later. I reach for the smallest thing that meets the requirement, not the largest that impresses, and let telemetry tell me when to revisit one. The habit came from presenting technical decisions to leadership and non-technical stakeholders, who were never going to accept "because it is best practice" as an answer.

Lately I have been pointing that at the infrastructure generative AI needs to run in production, in Go, Python, TypeScript, and Terraform on AWS: a retrieval augmented generation (RAG) service measured against a labeled evaluation set, a gateway that prices every inference call and fails over between providers, and a training pipeline where bad data cannot merge and no model reaches production without a human reading the evaluation. I am moving toward agentic RAG and document extraction. I am an AWS Certified Developer, studied computer science at Vanderbilt University and cloud engineering through the AWS Cloud Institute.

Personal

Most of what I do outside work comes down to getting people in the same room. That looks like cooking dinner for a friend group, or pulling friends who live in different cities onto a call to play through whatever strange indie game somebody turned up on Steam. The activity is mostly a pretext. The part I actually care about is that everyone showed up.

I cook a lot, and I recently restored my first cast-iron skillet, which took considerably longer than I expected. Mostly sanding rust off and putting thin coats of oil on it in the oven, over and over, until it stopped being orange. Worth it, though. Cast iron holds heat the way a thin pan cannot, so food hits the surface and actually sears the way you see in cooking videos.

I am in Nashville, so on a given week you might find me at a community event at my local bodybuilding gym, at a developer meetup at Bitcoin Park, or riding my road bike through a state park with no particular destination in mind.

I like learning for its own sake, and I have never really stopped. At the moment that means grinding through the long path of Math Academy's Mathematics for Machine Learning track, keeping up with deeplearning.ai, and working through DataCamp's AI and machine learning courses, then going and building something with whatever I just picked up. That last part is the honest explanation for most of the projects on this page. They started as things I wanted to understand rather than things I needed to ship.

Experience

Professional Experience

Roles, and the work that mattered in each.

AI Document Insights & Data Extraction Extern

Pfizer / Extern

Remote

  • Building an AI document extraction prototype replacing manual review of non-GMP clinical supply documents: scanned files parsed in Python with PyMuPDF, OpenCV, and OCR (Tesseract, PaddleOCR) to extract key fields and classify documents.
  • Implementing retrieval over the extracted corpus with Retrieval Augmented Generation (RAG) pipelines, metadata filtering, LangGraph, LlamaIndex, and a vector database.
  • Deploying a searchable interface for querying the extracted document corpus.

Cloud Engineering Resident

AWS Cloud Institute

Remote

  • Trained in cloud application development through AWS Cloud Institute, an official AWS program built and taught by AWS engineers.
  • Built event-driven workflows with AWS Lambda, Amazon API Gateway, Amazon DynamoDB, Amazon S3, Amazon SNS, and Amazon SQS, orchestrating multi-step processes with AWS Step Functions, and deploying containerized applications with Docker, Amazon ECS, Amazon EKS, and Amazon ECR.
  • Built LLM applications on Amazon Bedrock with LangChain, chaining model calls and managing conversation memory, and grounded them in domain data through Retrieval Augmented Generation (RAG) over Amazon Bedrock Knowledge Bases, applying prompt engineering and configuring Amazon Bedrock Guardrails for responsible AI.
  • Applied Amazon Rekognition and Amazon Textract for image analysis and document extraction, practicing Infrastructure as Code (IaC) and CI/CD, with least-privilege IAM and observability through Amazon CloudWatch and AWS X-Ray.
  • Provisioned infrastructure as code in CloudFormation, AWS CDK, and Terraform, and built CI/CD pipelines with AWS CodeBuild and CodePipeline across DevOps and DevSecOps workflows.

Business Systems Analyst

Price Printing

Nashville, TN

  • Built and maintained ETL pipelines standardizing multi-channel client-provided mailing address data to USPS formatting standards and verifying deliverability before print, automating ingestion in Python for batches of up to 250,000 addresses, where client-typical 8–15% defect rates meant reprints, wasted postage, and delays cascading into other clients' jobs.
  • Modeled reporting data dimensionally with star and snowflake schemas, surfaced through SQL and BI dashboards.
  • Integrated CRM, order management, and inventory systems, resolving the cross-platform data inconsistencies that were breaking fulfillment workflows.
  • Presented operational findings to leadership and worked with non-technical stakeholders to turn reporting into decisions.

Freelance Developer

Independent

Nashville, TN

  • Engineered automated ETL workflows to ingest, normalize, and validate raw client data from disparate third-party sources, enforcing the strict data hygiene and schema consistency required for trusted downstream workflows.
  • Delivered short-cycle contract engagements integrating client systems with industry-specific CRMs and third-party APIs, working against vendor platforms.
  • Built serverless applications on AWS for client workloads in React and Node.js.

Software Engineering Resident

Codesmith

Remote

  • Designed and built RESTful APIs and microservice architectures, applying async programming and data modeling across relational and non-relational databases.
  • Completed an intensive full-stack curriculum in JavaScript, TypeScript, Node.js, React, and Express, covering backend engineering, system design, and DevOps.
  • Worked in Agile cycles with pair programming, code review, and iterative delivery, building and contributing to open source software alongside a team.

Infrastructure and Monitoring Analyst

UPS

Nashville, TN

  • Contributed to an on-premise to hybrid AWS migration, working hands-on with Lambda, API Gateway, S3, RDS, DynamoDB, SNS/SQS, and CloudWatch across backend infrastructure supporting facility operations.
  • Developed incident response and monitoring for an 80,000-package sort, one of four daily sorts, setting CloudWatch thresholds against a 15–20 minute downtime target where faults drove missorts and service failures.
  • Transformed post-mortem incident reports into dimensional OLAP models that isolated analytical queries from transactional traffic, surfacing the recurring faults behind extended downtime so they could be addressed before repeating.
  • Decoupled facility systems with SNS and SQS event-driven workflows to route operational alerts across the distribution environment.
  • Diagnosed and resolved PLC faults on automated conveyor and sorting systems to protect throughput targets.

Digital Marketing Analyst

Independent

Nashville, TN

  • Consolidated campaign data from disparate ad platforms and Google Analytics sources with SQL and Excel models, normalizing inconsistent exports into one dataset that reporting and attribution depended on.
  • Analyzed campaign performance in Python, SQL, and Power BI, surfacing conversion rate, ROAS, and audience engagement trends that drove spend and targeting decisions.
  • Ran direct-to-consumer paid campaigns across multiple e-commerce accounts, owning budget allocation, audience targeting, and creative iteration.

Education

Education

Academic background and certifications.

AWS Cloud Institute

Cloud Engineering

Codesmith

Software Engineering Immersive

Vanderbilt University

Computer Science

Certifications

  • AWS Certified Developer
  • AWS Certified Cloud Practitioner
  • AI Engineer for Data Scientists Associate (DataCamp Exam)
  • Data Engineer Associate (DataCamp Exam)
  • Data Science with AI – Professional Certification (Codecademy 12 week training)
  • Apollo Graph Developer

Currently studying and developing for the AWS Certified Generative AI Developer – Professional certification.

Projects

Side Projects

A few cool things I've recently studied and implemented.

retrain-pipeline

Training and governance

Retraining pipelines automate the training but not the judgment: no contract on incoming labels, no lineage from a model back to the rows that made it, and one bad batch can promote itself to production.

CI-driven MLOps retraining pipeline on Amazon SageMaker with no endpoint, NAT gateway, or GPU, so idle cost is cents of S3 storage: Great Expectations gates new labeled data as a required check on every pull request, and DVC versions each dataset in an S3 remote with git holding only the content-hash pointer, which carries into the training job name so a resubmission of the same data is rejected as a duplicate.

A custom Go CLI available to the CI runner submits the job and registers the result as PendingManualApproval, under four minutes from merge, using IAM roles scoped by ARN with no wildcard actions that omit UpdateModelPackage, so no model is promoted without a human and CI could not approve one if it tried.

Stack: Go · Python · Amazon SageMaker · Model Registry · Great Expectations · Data Quality Gates · DVC · Dataset Lineage · Human-in-the-Loop Approval · Idempotent Job Submission · Least-Privilege IAM · scikit-learn · S3 · Terraform · GitHub Actions OIDC · CloudWatch

inference-gateway

Serving

A raw model endpoint has no per-caller identity, no throttling, no failover, and no answer to "what is p95 right now," so every team calling it rebuilds the same retry loop and nobody can attribute the bill.

LLM inference gateway in Go fronting Amazon Bedrock: Server-Sent Events (SSE) token streaming, per-key authentication and rate limiting, per-request cost metering, and multi-provider failover behind per-provider circuit breakers.

Per-key token buckets decide throttling in middleware, so a rejected request never reaches a paid API, and every response is priced by the model that actually answered. Roughly 23 microseconds of serving overhead per request single-threaded, from an 8.6 MB distroless image. Prometheus and Grafana observability; deployed on Amazon ECS Fargate with Terraform.

Stack: Go · Amazon Bedrock · Server-Sent Events · Rate Limiting · Cost Metering · Circuit Breakers · Multi-Provider Failover
ECS Express Mode on Fargate · Terraform · Docker · GitHub OIDC · Prometheus · Grafana · CloudWatch · CloudFront · TypeScript · React

rag-api

Retrieval

A RAG service that has never been measured is a demo. Nobody can separate a retrieval miss from a generation error, and a confidently wrong answer reads exactly like a right one.

Go service on Amazon ECS Fargate, two endpoints. /ingest splits a document on its heading structure at 800 runes and embeds each chunk with Titan v2 into pgvector. /query pulls 20 candidates by cosine distance, reranks them to 5 with Cohere Rerank, and returns { answer, sources[] } so a wrong answer traces to whether retrieval or generation failed.

Retrieval changes are measured against 35 hand-labeled questions, not argued. An offline harness took passage recall@5 from 57.1% to 77.1% across four changes, each measured before it was adopted. Hybrid search looked obvious and lost, so it is not in the service. A separate model grades faithfulness, validated against deliberately corrupted answers first to prove it can fail one: 70 graded responses, zero hallucinated claims, including 8 where retrieval returned nothing and the model declined instead of inventing.

Stack: Go · Amazon Bedrock · Titan v2 Embeddings · Cohere Rerank · pgvector · Source Citations · Evaluation Harness · LLM-as-Judge · RDS PostgreSQL · ECS Express Mode on Fargate · Multi-Tier VPC · Terraform · Docker · Distroless · GitHub OIDC · Secrets Manager · CloudWatch

christiansantiago.dev

Edge and delivery

Static hosting hides its own risks: a bucket left public, a deploy nobody can reproduce, and long-lived keys sitting in CI forever. This one has none of the three.

A Cloud Resume Challenge build. Static page in a private S3 bucket reachable only through CloudFront by Origin Access Control, so the bucket has no public URL at any point. The visitor counter runs on a Go Lambda, and increments through a single atomic DynamoDB update rather than a read followed by a write, so concurrent visitors cannot race each other into a lost count.

Every resource is Terraform against remote state and every deploy is a git push authenticated by GitHub OIDC, so no long-lived AWS keys exist anywhere, including in CI. The DynamoDB client sits behind a Go interface, so the handler unit-tests against a fake with no AWS calls at all, and a Playwright check runs against production after each deploy: if the counter does not render on the live page, the deploy failed.

Stack: Go on Lambda arm64 · DynamoDB · Atomic Increments · API Gateway HTTP API · S3 · CloudFront · Origin Access Control · Route 53 · ACM · Terraform · Remote State · GitHub OIDC · Keyless Deploys · Playwright E2E · CloudWatch

Resume

Resume

Highlights here, full PDF linked below.

Highlights

My work has consistently sat where cloud architecture meets business intelligence: close enough to the infrastructure to build it, close enough to the data to know what it is for.

Generative AI infrastructure

Three services deployed and verified on AWS covering retrieval, inference serving, and training governance: retrieval tuned against a labeled evaluation set rather than argued about, per-request token and cost accounting on a streaming gateway, data quality gates that run at the pull request, and model promotion that a human has to sign.

Cloud and platform engineering

Experience building AWS-hosted services in Go, Python, and TypeScript, REST APIs, Terraform defined infrastructure, and event-driven workflows on SNS and SQS, deployed through CI that authenticates by OIDC rather than stored keys.

Monitoring and incident response

Track record of shortening incident response by setting CloudWatch alerting thresholds, routing operational alerts through event-driven workflows, and feeding incident telemetry into the reporting pipelines behind failure analysis.

Data engineering and reporting

Background in ETL pipelines for client-supplied feeds, star and snowflake dimensional models that separate analytical reads from the transactional systems they compete with, and SQL-backed reporting dashboards presented directly to leadership.

The PDF

My current resume is viewable in the browser or available to download directly below.

View Resume Download PDF

Contact

Connect

Email, GitHub, and LinkedIn.

Always open to connecting, especially about generative AI infrastructure on AWS: retrieval, inference serving, and the evaluation and guardrail work that decides whether either is safe to put in front of users.