Rui Jin
Current focus
Current focus
Rec· Since Mar 2025
AI training & LLM evaluation
I work on the human side of AI training. On Outlier and AfterQuery, platforms that AI companies use to train and evaluate their models, I write prompts, grade model answers, and help improve the data large language models learn from.
Step 01
Write prompts
Write and refine prompts across classification, summarization, code generation, and multi-step reasoning, including hard ones built to expose where models fail.
Step 02
Grade responses
Score and compare responses against detailed rubrics for accuracy, reasoning, and instruction following, with a justification for every rating.
Step 03
Fix hallucinations
Flag hallucinations and factual errors, and rewrite weak answers into clear, correct references.
Step 04
Review code
Check generated code for correctness and readability, using 7+ years of engineering experience.
Evidence
Evidence
7+
Years building ML and software
Outlier & AfterQuery · Nex · Upwork · Customized Limited · Cloudbreakr
20+
Full-stack and AI projects for clients
Shipped on Upwork with a 5-star client rating
62%
Test videos counted within one rep
My on-device rep counter on the RepCount-A benchmark · best published pose-based method: 56%
0.57
Correlation with human rubric grades
Qwen3.5 4B judging locally on a 4 GB GPU, 320 BiGGen-Bench answers · GPT-4 as judge: 0.56
About
Full-stack first, then real-time vision. Now I help train and evaluate LLMs.
I started as a full-stack developer and moved into AI/ML because it's the fastest-growing part of software, and I wanted to build my career there. Most of my work now sits where machine learning meets real products. I train and tune models in Python with TensorFlow and scikit-learn, and I'm just as comfortable building the React or Next.js app that puts a model in front of users, or the AWS and Google Cloud setup it runs on.
At Nex, I worked on real-time body tracking for motion games on iOS and Android phones. Tracking had to run at 30 frames per second on the phone itself, and two of the hardest problems were encoding high-quality video and keeping the network connection stable. I trained and optimized lightweight vision models and built the data pipelines and debugging tools we used to test and ship them.
I also built a marine object detection system for a marine company: a YOLO26 model, stable enough to run on phones and drones, that spots kelp, plastic bottles, surfers, whales, boats, and containers. Before Nex, I delivered 20+ full-stack and AI integration projects on Upwork.
Today I do LLM training and evaluation work with Outlier and AfterQuery, and in May 2026 I passed micro1's AI interview to become a certified AI/ML Engineer. Next, I'm looking for an AI/ML Software Engineer role on an AI-focused team, where I can build models and the software that ships them for the long run.
At a glance
- Currently
- AI training & LLM evaluation
- Platforms
- Outlier, AfterQuery
- Looking for
- AI/ML Software Engineer roles
- Based in
- Surigao del Norte, Philippines
- Certified
- Certified AI/ML Engineer, micro1
- Main tech
- Python · TensorFlow · OpenCV · YOLO · LLMs · React · Next.js · AWS · Google Cloud
What I work on
What I work on.
01 / Focus
LLMs & NLP
Prompt engineering, rubric-based evaluation, hallucination review, and LLM features in client products
02 / Focus
Computer vision
Deep learning models for real-time object detection, tracking, and pose estimation, from body tracking on phones to marine detection on drones with YOLO26
03 / Focus
On-device ML
Making models smaller and faster so real-time tracking runs at 30 fps on iOS and Android phones, plus the camera, video encoding, and inference code around them
04 / Focus
MLOps
Data pipelines, training and evaluation workflows, CI/CD, deployment, and debugging tools that check accuracy in real-world conditions
05 / Focus
Full-stack
React, Next.js, Flask, and Node.js apps and REST APIs on AWS and Google Cloud that put models in front of users
06 / Focus
Data
ETL pipelines, NoSQL data stores, analytics, and client-facing dashboards
Experience
Seven years of shipping.
Mar 2025 – Present
Outlier & AfterQuery
AI Training & Evaluation Specialist
Help AI companies train and evaluate their large language models (LLMs) through the Outlier and AfterQuery platforms, improving the data those models learn from and are measured against.
- Write and refine prompts across classification, summarization, code generation, and multi-step reasoning tasks, including hard prompts built to expose where models fail.
- Score and compare model responses side by side against detailed rubrics for accuracy, reasoning, and instruction following, and write a clear justification for each rating.
- Flag hallucinations and factual errors, and rewrite weak answers into clear, correct reference responses.
- Review generated code for correctness and readability, drawing on 7+ years of software and ML engineering.
LLMs · Prompt engineering · LLM evaluation · RLHF · Code review
Dec 2022 – Mar 2025
AI/ML Engineer
Worked on the real-time body tracking behind Nex's motion-controlled games, where players control the game by moving in front of their phone's camera.
- Built and optimized lightweight computer vision models for real-time pose and gesture tracking on iOS and Android phones, targeting 30 fps with all inference running on-device.
- Worked on two of the hardest parts of the pipeline: encoding high-quality video on the phone and keeping the network connection stable.
- Wrote low-latency code for camera input, sensor fusion, and model inference, and connected it to the SDK our game designers and outside Unity developers used to build motion-controlled games.
- Set up data pipelines and MLOps tooling to train, evaluate, and deploy models, plus debugging tools to check tracking accuracy across different lighting and room setups.
Computer vision · Pose estimation · On-device ML · Sensor fusion · MLOps
Jun 2021 – Nov 2022
Upwork
Freelance Software Developer
- Completed 20+ full-stack and AI integration projects while keeping a 5-star client rating.
- Built and shipped web apps with React, Next.js, Flask, and Python on AWS and Google Cloud for SaaS, e-commerce, and analytics clients.
- Added LLM features and third-party API integrations to client products, and automated manual workflows along the way.
React · Next.js · Flask · Python · JavaScript · AWS · GCP
Sep 2019 – Nov 2022
Customized Limited
Full-Stack Developer
- Built, tested, and deployed front-end and back-end features for the company's platform using Java, Python, and Node.js.
- Used Python to connect the platform to the company's computer vision and machine learning models.
- Worked with full-time engineers and ML interns who processed and labeled the video data behind those models.
Java · Python · Node.js
Jun 2018 – Aug 2019
Cloudbreakr
Hong Kong
Software Engineer Intern
- Built features for an influencer marketing platform, using PHP and Laravel on the back end and JavaScript and React on the front end.
- Added tracking charts to influencer profiles and built dashboards that showed brand clients the ROI of their campaigns.
PHP · Laravel · JavaScript · React
Projects
Selected projects.
01 / Featured project
On-Device Rep Counter
Counts exercise reps live from a camera, right in the browser: a small PyTorch model, trained from scratch, reads MediaPipe body landmarks and runs on INT8 weights. On the RepCount-A benchmark it counts 62% of test videos within one rep; the best published pose-based method counts 56%.
- Within one rep
- 62%
- Model
- 106 KB
- Per frame
- 0.35 ms
PyTorch · MediaPipe · ONNX · INT8 quantization · TypeScript
02 / Featured project
Small LLM Judges vs. Human Rubric Grades
Can a small open model, run on a 4 GB GPU, grade answers against detailed rubrics like a person? On 320 human-scored BiGGen-Bench answers, Qwen3.5 4B (4-bit) matched the human grades as closely as GPT-4 and Claude 3 Opus did. It graded harsher than people; a cross-validated calibration brought it within one point on 86% of answers.
- Pearson r
- 0.57
- GPT-4 judge
- 0.56
- Within 1 point
- 86%
Python · llama.cpp · LLM-as-a-judge · Rubric evaluation
03 / Project
EmbeddingGemma 2 Code Search Benchmark
An open benchmark of Google's EmbeddingGemma 2 against the first EmbeddingGemma on code search, at 768, 512, 256, and 128 dimensions, plus CPU speed and memory. At 256 dimensions, EmbeddingGemma 2 matched the original model at 768.
Python · PyTorch · sentence-transformers · Hugging Face
04 / Project
Marine Object Detection
Client work · NDA
Real-time detection of marine objects for a marine company: kelp, plastic bottles, surfers, whales, boats, and containers. Built on YOLO26 as a stable model that runs on phones and drones.
- Input
- Video from phones and drones
- Model
- YOLO26 object detection
- Finds
- Kelp, plastic bottles, surfers, whales, boats, containers
- Runs on
- Phones and drones, in real time
Python · YOLO26 · TensorFlow · OpenCV · Drone SDK
05 / Project
Instagram Avatar Protection
Client work · NDA
Uses image recognition and watermarking to catch and prevent unauthorized reuse of users' Instagram profile photos.
- Input
- Users' Instagram profile photos
- Method
- Image recognition and watermarking
- Result
- Catches and prevents unauthorized reuse
Python · TensorFlow · OpenCV · Instagram API
06 / Project
SaaS Landing Page
A fast, responsive landing page for a SaaS product, with email integration through Resend.
Next.js · React · Tailwind CSS · Resend
Writing & Talks
Writing & Talks
EmbeddingGemma 2 at 256 Dimensions Matches the Original at 768
Benchmark · Medium
Skills
The stack behind the work.
- 01 / Languages
- Python
- TypeScript
- JavaScript
- Java
- C#
- SQL
- PHP
- 02 / Machine learning & computer vision
- PyTorch
- TensorFlow
- scikit-learn
- OpenCV
- YOLO
- Deep learning
- Computer vision
- Object detection
- Pose estimation
- Model evaluation
- 03 / On-device & real-time ML
- On-device ML
- Model optimization
- INT8 quantization
- ONNX
- Real-time inference
- Sensor fusion
- Video encoding
- iOS
- Android
- 04 / LLMs & NLP
- LLM evaluation
- Prompt engineering
- Rubric-based evaluation
- Hallucination review
- LLM integration
- Natural language processing (NLP)
- 05 / Data & MLOps
- ML pipelines
- MLOps
- ETL pipelines
- NoSQL
- Data analytics
- Data visualization
- 06 / Cloud & DevOps
- AWS
- Google Cloud (GCP)
- Docker
- Kubernetes
- CI/CD
- Git
- 07 / Backend & APIs
- Node.js
- Flask
- Laravel
- REST APIs
- 08 / Frontend
- React
- Next.js
- Tailwind CSS
Education
Education
Bachelor of Computer Science
- Graduated with honors, focusing on software development and AI
- Thesis: AI-driven web applications
Certifications
Certifications
May 2026
Certified AI/ML Engineer
micro1
Fun facts
Off the clock.
Note 01
Powered by espresso
My models run on GPUs. I run on espresso.
Note 02
Still waiting for my Hogwarts letter
Harry Potter is my favorite book. Until the owl shows up, I'll keep casting spells in Python.
Note 03
Certified comedy addict
Zootopia, Ne Zha, and Despicable Me are my all-time favorite comedies. The Minions get me every single time.
Lines I live by
“I am the master of my fate, I am the captain of my soul.”
William Ernest Henley, Invictus “Stay hungry, stay foolish.”
Steve Jobs
Contact
Bring me the model that has to run in real time.
I'm open to AI/ML Software Engineer roles. Email is the quickest way to reach me, or you can book a short call.