Tran Quang Khai

Machine Learning / Computer Vision

With a strong foundation in Machine Learning, Deep Learning, and Computer Vision, I aim to pursue graduate studies to deepen my theoretical understanding and learn to deploy and maintain machine learning systems in production. I am particularly interested in building automated data and machine learning pipelines.

Summary

I focus on building practical AI systems with an emphasis on computer vision, model deployment, and production-oriented pipelines. My recent work spans aesthetic-aware image cropping, real-time detection systems, and ML-driven automation across research and applied engineering settings.

Computer Vision Model Deployment ML Pipelines

Education

University of Information Technology VNUHCM

Bachelor of Science, Computer Science

Sep 2021 - Jun 2025
  • Current GPA: 8.43/10
  • Coursework: Artificial Intelligence, Computer Vision, Deep Learning, Data Mining

Experience

Lab Assistant

Multimedia Communications Laboratory

May 2024 - January 2026
  • Conduct research on aesthetic-aware image cropping using Large Language Models such as GPT and Gemini, contributing to a forthcoming publication in 2025.
  • Developed and published two datasets for aesthetic-aware image cropping, evaluated on those datasets to establish a benchmark.

Technical Specialist

Vietnam-Singapore Industry 4.0 Innovation Center, Eastern International University

Oct 2025 - Present
  • Participated in developing AI for HVAC systems to control the school's AC, achieving fan speed accuracy of 91%.
  • Integrated the YOLO11x model for people counting on the CCTV server for building management using DeepStream, achieving real-time human detection on multiple cameras.

Projects

Aesthetic-Aware Image Cropping

May 2024 - Present
  • Project description: Build a pipeline for prompt-based image cropping using aesthetic-awareness.
  • Methods:
    • Generative AI models like GPT and Stable Diffusion to generate synthetic datasets.
    • Leveraged models such as CACNet, SAM2, and Grounding DINO for image cropping.
  • Results:
    • Published two datasets: 1,000 manually labeled images and 4,000 automatically labeled.
    • mDSC: 0.522 using Gemini-2.0-Flash, and 0.453 using GPT-4o mini.
  • Tools: Python, GPT, Stable Diffusion, TensorFlow, Gemini, SAM2, Grounding DINO.

Building a Real-Time Human Detection System

Oct 2025 - Present
  • Project description: Using Nvidia DeepStream for real-time human detection on CCTV.
  • Methods:
    • Fine-tuning YOLO11x for maximum accuracy on human detection.
    • Converted the model into TensorRT to boost FPS and apply batch processing.
    • Integrated MediaMTX to transcode RTSP into low-latency WebRTC.
  • Results:
    • Optimized YOLO11x inference using TensorRT FP16 precision, enabling the system to scale from 1 to 10 concurrent video feeds on a single GPU while maintaining 20 FPS real-time.
  • Tools: Python, DeepStream, Mediamtx, Caddy.

Skills and Technologies

Languages: C++, Python.
AI & Computer Vision: NVIDIA DeepStream, TensorRT, PyTorch, OpenCV, Ultralytics YOLO, CUDA.
DevOps & Deployment: Docker, Docker Compose, Git, Linux (Ubuntu/WSL).
Cloud Platforms: Azure (familiar with compute and basic networking).
IELTS: 8.0 (L: 8.5, R: 9.0, W: 7.0, S: 6.5).