Skip to content
View vilyam's full-sized avatar

Block or report vilyam

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
vilyam/README.md
header taglines

LinkedIn Email Upwork GCP Pro Location Status Views

15+ years building production data and cloud systems for Walmart × Google, Tiffany & Co., Procore, and ITV.

What I deliver: cloud-native data and AI platforms across GCP, AWS, Azure, and self-hosted open source — warehouses, lakehouses, Delta and Iceberg, medallion ELT/ETL, governed Master Data Management (MDM) — for AI, marketing, sub-second real-time streaming, and terabyte-scale batch workloads. Serverless batch with 0-cost-when-idle economics. Retrieval-Augmented Generation (RAG) + Model Context Protocol (MCP) retrieval pipelines for AI products on Vertex AI.

Architect-level stewardship across the lifecycle: data governance & lineage, data quality (DQ assertions, schema contracts, observability), security & compliance posture (SOC2 · GDPR · HIPAA readiness), FinOps cost engineering, SLO-driven reliability, and pre-sale discovery and scoping. Translate business KPIs (ROI · ARPU · LTV) into measurable engineering outcomes.

Recent outcomes: $M+ Dataproc/Spark cost cut at Walmart × Google · >6× GCP cost cut at ogment.ai · 3 TB → Elasticsearch in 3 h for <$100 · CI/CD 2 h → 30 min at Procore · 30+ enterprise sources unified at Tiffany & Co.

Open to engagements as Data Architect, Lead / Principal Data Engineer, Fractional CTO, Solution Architect, or AI Strategy lead. Streaming (Kafka · Flink · Pub/Sub · Dataflow), batch (Spark · Hadoop · Hive · Airflow · Dataproc · EMR · BigQuery), lakehouses (Delta · Iceberg), AI (RAG · MCP · Vertex AI) — on GCP, AWS, Azure, or self-hosted open source. Solo or team lead.

Tech DNA

Cloud GCP AWS Azure Cloudflare
AI / ML Vertex AI Document AI MCP Gemini Claude OpenAI pgvector
Streaming Kafka Kafka Streams Flink Beam Pub/Sub Dataflow Redis Streams RabbitMQ
Batch Spark Hadoop Hive Dataproc EMR Airflow Cloud Composer MWAA Dagster dbt Dataform Databricks Fivetran Airbyte AWS Glue Step Functions DuckDB
Warehouse BigQuery Redshift Athena Cloud Spanner ClickHouse Postgres Redis Elasticsearch MongoDB Cassandra ScyllaDB Iceberg TimescaleDB Firestore MySQL Oracle GCS S3 Cloudflare R2
Backend Spring Boot Play Framework Akka Micronaut FastAPI Django Flask NestJS Netty
Languages SQL Java Scala Python TypeScript Node.js Bash
Platform Terraform Kubernetes Docker GitHub Actions CircleCI Jenkins Cloud Build Firebase Supabase
Principles KISS / Occam's Razor DRY SOLID YAGNI WORA Let it crash
Patterns Medallion Kimball Strangler-fig ES+CQRS Actor Model Event Choreography Hub-and-Spoke Landing Zone GitHub Flow Trunk-based Development
Governance SOC2 GDPR HIPAA Data Contracts Data Lineage Data Quality FinOps OpenTelemetry SLOs

Current Engagement  Active

B2B Master Data Management (MDM) platform — Solution Architect & Data Lead · 2025–present · NDA · author of the key architectural decisions.

Reframed the scope: a legacy ETL replatform was actually a Master Data Management problem. 5 consumer systems each pulling their own version of the truth, no governance, brittle to change. Migrated to a governed MDM hub — medallion ELT, Kimball dimensional silver, strangler-fig cutover, reverse-ETL fan-out. Live in production, no downtime, no incident.

MDM Master Data Hub Medallion ELT Kimball Dimensional Reverse-ETL Strangler-fig Hash-diff parity SQL Dagster Dataform BigQuery Terraform

Solo Build  Live

nexalgames.com — live-video arcade platform · 2024–present · solo end-to-end.

A multi-region GCP platform that streams live video of physical claw machines to players worldwide, with a real-money token economy. The combination — live video, hardware, payments, multi-region — usually needs an 8–12 person team for 18 months. One operator owning it all: architecture, infrastructure, hardware integration, video pipeline, payment safety. Live in production.

GCP Landing Zone Multi-region Live Video Hardware Integration Real-money Idempotency Event Choreography Java Terraform

Shipped

Greenfield AWS data warehouse · 2026 · NDA · AWS-native

Designed and built the data warehouse end-to-end at 250M rows/month — daily Parquet ingestion, DuckDB for local batch exploration, Elasticsearch search layer, sub-second query response. Stack: Redshift · Athena · Glue · Step Functions · Lambda · S3 · DuckDB.

AWS Redshift Athena SQL Glue Step Functions Lambda S3 DuckDB Elasticsearch Parquet

ogment.ai · 2025

Solo 6-hat platform lead (Cloud Architect · Data Architect · Data Engineer · Data Analyst · DevOps · Lead SWE). >6× GCP cost cut. End-to-end DWH in 3 weeks (industry baseline: 6 months). BigQuery → Elasticsearch at 3 TB / 3 h / <$100 — called a record on design review by Elastic SMEs. Built RAG pipelines and MCP-driven agent flows over the candidate corpus — hybrid retrieval, reranking, multi-vendor LLM evaluation (5 dataset providers benchmarked pre-purchase).

RAG MCP Hybrid Retrieval Reranking LLM Evals Landing Zone Elasticsearch BigQuery SQL Supabase Python Bash Terraform GCP

swiftbuild · 2024–2025 · zoning & document intelligence

Document AI + Vertex AI extraction pipelines: PDF → structured zoning rules → geospatial search. Structured generation with schema-validated LLM output. Org-level GCP infrastructure (Shared VPC, Cloud SQL with IAM connector, Artifact Registry, CDN + load balancing) designed to SOC2-readiness (least-privilege IAM, audit logs, secrets rotation, encryption at rest + in transit).

Document AI Vertex AI Structured Generation LLM Extraction Landing Zone Cloud SQL Geospatial Cloud Functions Terraform

Procore · 2024 · SoftServe

CI/CD pipeline rebuilt: 2 h → 30 min. Databricks pipeline optimization with CDC (Change Data Capture) ingestion via Debezium + Airbyte. Datadog observability rework.

Databricks Delta Lake CDC Debezium Airbyte Kafka Datadog CircleCI

Tiffany & Co. · 2023–2024 · SoftServe

Tech Lead of the Data Integration & Ingestion team. Founding data engineer — unified 30+ enterprise sources into one DWH across GCP + Azure + Redshift.

GCP Azure Redshift Dataform SQL Fivetran Data Integration Team Leadership

Walmart × Google · 2020–2023 · SoftServe

Lead on Dataproc product. Built Dataproc-on-Demand and Dataproc-aaS. $M+ Dataproc/Spark cost cut for internal Walmart customers. Automic → Airflow and db2 z/OS → BigQuery migrators.

Dataproc Airflow BigQuery SQL Java Scala Python Bash Kubernetes

ITV · 2020 · SoftServe

Production AWS → GCP migration with zero downtime. Apache Beam + Airflow + BigQuery, Terraform-coded end to end.

AWS GCP Apache Beam Airflow Terraform

Earlier · 2008–2020

Tech Lead at talentreef — built and scaled the team 0 → 12, delivered the full platform rebuild end-to-end (Scala · Kafka · Cassandra · ES+CQRS · microservices). Senior at Lotusflare (Scala · Spark Streams · Kafka Streams · Redis · Cassandra at scale). Lead at Gamestars (Akka · Play · Actor Model — LoL-Elo matchmaking, MVP in 3 weeks). Delphi · .NET · Java from 2008.

Scala Akka Spark Streams Kafka Streams Cassandra Redis ES+CQRS Microservices Actor Model Team Leadership

Work with me

What I take on: legacy ETL replatforms · runaway cloud bills · AI / RAG retrofits into systems not built for AI · greenfield data + AI platforms delivered in weeks. Fractional or end-to-end — solo or alongside your team.

Email · LinkedIn · Upwork — Expert-Vetted · GCP Pro Cert

footer

Architecting production systems since 2008 — Kyiv, Ukraine.

Popular repositories Loading

  1. similar-html-element-finder similar-html-element-finder Public

    Scala

  2. ads-txt-crawler ads-txt-crawler Public

    Scala

  3. davio-parsing-tool davio-parsing-tool Public

    Parsing tool to read the .doc/.docx files and extract the valuable data, transform it and load as a CSV.

    Python

  4. devops-tools devops-tools Public

    Shell

  5. elasticsearch-single-node-tf-module elasticsearch-single-node-tf-module Public

    Elasticsearch Single node Terraform module with proper config management and Kibana instance

    HCL

  6. vilyam vilyam Public