HungryNeko avatar HungryNeko HungryNeko's Blog, 又是一条咸鱼呢...

HungryNeko profile decoration HungryNeko profile decoration HungryNeko profile decoration HungryNeko profile decoration

About me

  • @HungryNeko
  • Backend Engineering + AI Engineering (CV, speech, RL, multimodal)
  • Current focus: production AI pipelines (ONNX inference, Dockerized services, async cloud workflows, agent tools)
  • Interests: IoT backend, speech processing, AI agents, and research-to-production engineering
  • More technical notes: Blog Posts

Tech Stack

AreaStack
LanguagesPython, C/C++, Java, C#, SQL
AI/MLPyTorch, TensorFlow, Hugging Face Transformers, LangGraph, scikit-learn, OpenCV, YOLOv8, ONNX, ONNX Runtime
Speech & MultimodalWhisper, MossFormer2, SpeechBrain, WeSpeaker, cross-lingual speaker verification, code-switch analysis
Backend & CloudNode.js, Flask, REST APIs, JWT, MySQL, SQLite, Docker, Linux, MQTT, AWS (EC2/S3/SQS/Lambda/DynamoDB/IAM/Secrets Manager), Azure
FrontendReact, Angular, HTML
EngineeringGit, GitHub Actions, CI/CD, async job pipelines, PyQt, QT, TCP/IP, COLMAP, 3D Gaussian Splatting

Selected Projects

  • AI-Powered Rental Management Agent Platform
    • Tech: Flask, SQLite, JWT, LangGraph, MCP, SQL tool, Python tool, chart tool, PDF tool, vector RAG.
    • Features: decoupled MCP-based agent integration, natural-language data entry, record update/query, charting, PDF/file/image handling, automation, and retrieval-augmented contract/document search.
    • Github Agent Part (refining)
  • AMECxSV: Metadata-Driven Calibration for Cross-Lingual Speaker Verification
    • Tech: frozen speech encoders, metadata-aware score calibration, language/duration features, lightweight MLP.
    • Focus: cross-lingual speaker verification, multilingual trials, confidence-based abstention.
    • arxiv
  • GBC: Gaussian-Based Colorization and Super-Resolution for 3D Reconstruction
    • Tech: optical-flow super-resolution, temporal colorization, FFmpeg preprocessing, COLMAP + 3D Gaussian splatting.
    • Publication: ACM SIGGRAPH VRCAI 2024.
    • Blog Paper Demo Github
  • BMS^3: Bayesian Modeling Based SwinUNet Segmentation on Self-distillation Architecture
    • Tech: Bayesian modeling, SwinUNet backbone, self-distillation, cross-domain segmentation setup.
    • Publication: ICONIP 2025.
    • Blog Paper Reference List
  • Safety-driven Path Selection Using Reinforcement Learning in Autonomous Driving
    • Tech: Q-learning, dynamic confidence update, noisy-source filtering, OpenStreetMap-based routing context.
    • Publication: RSAE 2025.
    • Blog
  • Multilingual Speech Separation + Code-switch Correction Pipeline
    • Tech: MossFormer2, Whisper, PyTorch, SpeechBrain/WeSpeaker, custom TDNN/SincNet variants.
    • Experiments: short-window cross-lingual speaker verification benchmark across ECAPA, x-vector, WavLM, Resemblyzer, and custom models, with ablation + speed/accuracy comparison
  • AI Cloud Album (AWS)
    • Tech: Flask, JWT, S3, SQS, Lambda, DynamoDB, IAM, Secrets Manager, status-driven async workflow (uploaded -> processing -> complete/failed).
    • AI Deployment: YOLOv8 inference exported to ONNX and packaged with Docker for reproducible cloud inference.
    • Blog
  • R2 Gateway
    • Tech: Flask, Docker, Cloudflare R2, S3-compatible APIs, Flask-Limiter.
    • Features: Dockerized R2 gateway with token-based access control, public/private bucket policy, traffic and operation-quota guardrails, and health/usage endpoints.
    • Blog
  • SAR Data Pipeline with YOLOv8
    • Tech: YOLOv8, OpenCV preprocessing, augmentation pipeline, format conversion, classification/detection/OBB training.
    • Metrics: 35% effective data expansion, 12% accuracy improvement over baseline.

Note: Some projects are in private repositories (course/research or confidentiality reasons).