Data Engineer | AI Systems Builder | Distributed Systems Enthusiast
I build scalable data infrastructure for AI systems.
From ingestion pipelines to distributed computing and GPU-accelerated workloads, I focus on designing high-performance architectures that power intelligent applications.
Currently pursuing a BS in Artificial Intelligence (FAST-NUCES) with a strong emphasis on:
- Data Engineering
- Distributed Systems
- AI Infrastructure
- High-Performance Computing
- Designing scalable ETL pipelines for AI workloads
- Building distributed and microservice-based AI systems
- Exploring GPU-accelerated data processing (CUDA, RAPIDS, parallel computing)
- Studying large-scale RAG systems and multi-agent reasoning
- Strengthening cloud-native architecture skills (AWS, Azure, GCP)
- ETL pipeline design & optimization
- Data scraping & transformation
- Data curation for ML pipelines
- Data governance & quality assurance
- Microservices architecture
- Containerization (Docker, Kubernetes)
- CI/CD pipelines
- Cloud-native deployments
- NLP systems (BERT variants, embeddings, RAG)
- PyTorch, TensorFlow, Scikit-learn
- Multi-agent reasoning frameworks
- GPU-based acceleration & parallel computing
- C#, .NET, Java
- Python (FastAPI, Flask)
- Node.js, Express
- REST & gRPC APIs
Microservice-based NLP system that generates creative stories from a topic and converts them into narrated audio.
Tech Stack:
Python, FastAPI, gRPC, Databricks Dolly v2, edge-TTS, Gradio, Docker
Flask-based demonstration of quantum-resistant encryption using modern KEM algorithms.
Tech Stack:
Python, Flask, liboqs-python
Algorithms: ML-KEM-512, FrodoKEM, BIKE, McEliece
Designing and evaluating a retrieval-grounded multi-agent debate framework to analyze:
- Knowledge drift
- Self-consistency
- Evidence alignment
- Convergence behavior
Comparative evaluation against:
- Single-Agent RAG
- Self-Consistency
- Tree-of-Thoughts
Focus: Stability vs. drift under noisy retrieval pipelines.
- Optimized ETL pipelines for AI model training
- Led interns and distributed data engineering tasks
- Built structured failure management systems
- Delivered insights from large-scale scraped datasets
- Worked on Urdu Word Sense Disambiguation
- Experimented with mBERT, mDistilBERT, RobertaUrdu
- Contributed research for academic publication
Languages:
C#, Java, Python, C++, JavaScript, TypeScript
Data & ML:
PyTorch, TensorFlow, Scikit-learn, NLTK, SpaCy, NumPy, Pandas
Cloud & DevOps:
Azure, Docker, Kubernetes, CI/CD
Databases:
CosmosDB, MSSQL, NoSQL
Systems:
Unix, Shell, System Administration
- Architect large-scale AI data platforms
- Contribute to open-source distributed systems
- Build GPU-optimized AI infrastructure
- Advance real-time streaming & serverless architectures
- Develop reliable, scalable AI systems beyond experimentation
- π LinkedIn: https://www.linkedin.com/in/talha-syed-b49304219/
- βοΈ Blog: https://substack.com/@bitwisetitan
- π» GitHub: https://github.com/BitwiseTitan
- π§ Email: talha.syed1@outlook.com
Data is not just fuel for AI.
It is the architecture of intelligence.