About me
- @HungryNeko
- Backend Engineering + AI Engineering (CV, speech, RL, multimodal)
- Current focus: production AI pipelines (ONNX inference, Dockerized services, async cloud workflows, agent tools)
- Interests: IoT backend, speech processing, AI agents, and research-to-production engineering
- More technical notes: Blog Posts
Tech Stack
| Area | Stack |
|---|---|
| Languages | Python, C/C++, Java, C#, SQL |
| AI/ML | PyTorch, TensorFlow, Hugging Face Transformers, LangGraph, scikit-learn, OpenCV, YOLOv8, ONNX, ONNX Runtime |
| Speech & Multimodal | Whisper, MossFormer2, SpeechBrain, WeSpeaker, cross-lingual speaker verification, code-switch analysis |
| Backend & Cloud | Node.js, Flask, REST APIs, JWT, MySQL, SQLite, Docker, Linux, MQTT, AWS (EC2/S3/SQS/Lambda/DynamoDB/IAM/Secrets Manager), Azure |
| Frontend | React, Angular, HTML |
| Engineering | Git, GitHub Actions, CI/CD, async job pipelines, PyQt, QT, TCP/IP, COLMAP, 3D Gaussian Splatting |
Selected Projects
- AI-Powered Rental Management Agent Platform
- Tech: Flask, SQLite, JWT, LangGraph, MCP, SQL tool, Python tool, chart tool, PDF tool, vector RAG.
- Features: decoupled MCP-based agent integration, natural-language data entry, record update/query, charting, PDF/file/image handling, automation, and retrieval-augmented contract/document search.
- Github Agent Part (refining)
- AMECxSV: Metadata-Driven Calibration for Cross-Lingual Speaker Verification
- Tech: frozen speech encoders, metadata-aware score calibration, language/duration features, lightweight MLP.
- Focus: cross-lingual speaker verification, multilingual trials, confidence-based abstention.
- arxiv
- GBC: Gaussian-Based Colorization and Super-Resolution for 3D Reconstruction
- BMS^3: Bayesian Modeling Based SwinUNet Segmentation on Self-distillation Architecture
- Tech: Bayesian modeling, SwinUNet backbone, self-distillation, cross-domain segmentation setup.
- Publication: ICONIP 2025.
- Blog Paper Reference List
- Safety-driven Path Selection Using Reinforcement Learning in Autonomous Driving
- Tech: Q-learning, dynamic confidence update, noisy-source filtering, OpenStreetMap-based routing context.
- Publication: RSAE 2025.
- Blog
- Multilingual Speech Separation + Code-switch Correction Pipeline
- Tech: MossFormer2, Whisper, PyTorch, SpeechBrain/WeSpeaker, custom TDNN/SincNet variants.
- Experiments: short-window cross-lingual speaker verification benchmark across ECAPA, x-vector, WavLM, Resemblyzer, and custom models, with ablation + speed/accuracy comparison
- AI Cloud Album (AWS)
- Tech: Flask, JWT, S3, SQS, Lambda, DynamoDB, IAM, Secrets Manager, status-driven async workflow (
uploaded -> processing -> complete/failed). - AI Deployment: YOLOv8 inference exported to ONNX and packaged with Docker for reproducible cloud inference.
- Blog
- Tech: Flask, JWT, S3, SQS, Lambda, DynamoDB, IAM, Secrets Manager, status-driven async workflow (
- R2 Gateway
- Tech: Flask, Docker, Cloudflare R2, S3-compatible APIs, Flask-Limiter.
- Features: Dockerized R2 gateway with token-based access control, public/private bucket policy, traffic and operation-quota guardrails, and health/usage endpoints.
- Blog
- SAR Data Pipeline with YOLOv8
- Tech: YOLOv8, OpenCV preprocessing, augmentation pipeline, format conversion, classification/detection/OBB training.
- Metrics: 35% effective data expansion, 12% accuracy improvement over baseline.
Note: Some projects are in private repositories (course/research or confidentiality reasons).