LLaVA-OneVision-2.0 End-to-End Model Optimization

Turn Frontier AI intoBusiness Productivity

Built on our independently developed and open-source LLaVA-OneVision technology stack, Glint-VL-Video-8B provides vendor-native customization from model optimization and task adaptation to deployment for government and enterprise scenarios.

Try

MODEL TRAINING & REFINEMENT

Turn private data into proprietary model capability

The provides a full model-optimization stack spanning foundational training, parameter-efficient fine-tuning, and knowledge distillation, closed by objective evaluation and production deployment. Data stays on site while capability accumulates in the customer model.

  1. TRAINING

    Model Training

    Build datasets, training recipes, and version baselines around domain data and target tasks, with support for 4B and 8B VLM training.

    • Data governance and multimodal annotation
    • Training monitoring and checkpoint recovery
    • Model versions and experiment records
  2. FINE-TUNING

    Model Fine-Tuning

    Use SFT, LoRA, and QLoRA to adapt general models to customer tasks, domain knowledge, and output requirements.

    • Cost-efficient parameter adaptation
    • Private blind tests and A/B comparison
    • Controlled quality, cost, and delivery cycle
  3. DISTILLATION

    Model Distillation

    Transfer knowledge and business reasoning from larger models into smaller ones while balancing quality, latency, and private deployment cost.

    • Distillation for LLMs below 32B
    • Quantization, compilation, and inference tuning
    • Cloud, private, and edge deployment
  1. 01Data Assets
  2. 02Train / Tune / Distill
  3. 03Evaluate and Align
  4. 04Deploy and Deliver

GLINT-VL MODEL FAMILY

Multimodal perception for the physical world

Anchored by advanced video understanding and refined for security and banking environments, the Glint-VL family brings multimodal intelligence into real operational settings.

CODEC-STREAM VIDEO UNDERSTANDING

Vision-Language Model

Glint-VL-Video-8B is the first large video understanding model that uses codec streams as its visual unit, identifying key events, locating time ranges, and extracting evidence for video review, content retrieval, and process analysis.

  • Long-Video Understanding
  • Temporal Grounding
  • Event Retrieval
Video Understanding

SECURITY-SPECIALIZED MODEL

Security-Domain Vision-Language Model

Glint-VL-SE-8B is built on large-scale security imagery and instruction-tuned domain datasets, substantially improving recognition of critical security events and fine-grained person and vehicle attributes.

  • Critical Event Recognition
  • Person & Vehicle Attributes
  • Instruction Tuning
Object Discovery

BANKING PHYSICAL-WORLD MODEL

Glint-VL-BK-8B

V1.0

Banking Physical-World Vision-Language Model

Purpose-trained for branches, vaults, and office buildings, Glint-VL-BK-8B identifies human behavior, critical events, safety risks, and operational compliance issues for continuous, traceable analysis.

  • Bank Branches
  • Vault Security
  • Office Operations
  • Operational Compliance

PRODUCTS BUILT ON MODELS

Deliver model capability into operations

City management combines general video understanding with security-specialized models. In finance, general language models handle knowledge and documents while the banking model understands branches, vaults, and offices.

CITY

Visual Intelligence for City Management

The general video model understands long-running events and processes, while the security model recognizes governance events, people, vehicles, and fine-grained attributes.

Video-8BSE-8BCity Event Workflow
  • Road occupation, abandoned objects, and crowd events
  • Object trajectories, time ranges, and key evidence
  • Integration with existing city and video systems

FINANCE

Intelligent Finance

General LLMs and OCR process statements, financial reports, and due-diligence materials, while BK-8B handles physical-space security and operational compliance.

General LLMBK-8BIntelligent Finance
  • Assistants for statements, reports, and due diligence
  • Operational compliance across branches, vaults, and offices
  • Local 32B LLM, OCR, AI gateway, and GBOX

GLINT LAB MODEL ECOSYSTEM

Open research and proprietary models form one foundation

Research spans VLMs, visual foundations, multimodal embeddings, multimodal RAG, face recognition, and 3D vision.

VLM

Open Source

LLaVA-OneVision 2.0

统一图像视频的全帧率视觉语言大模型

  • 多模态大模型

VLM

Open Source

LLaVA-OneVision-1.5

全开源的多模态大语言模型训练框架,

  • 多模态大模型

VLM

AAAI 2026

ViCToR

利用视觉token重建提升多模态大模型的方法、

  • 多模态大语言模型

VLM

CVPR 2026

StreamingRVOS

基于大语言模型的流式参考视频分割方法

  • 流式视频分割
  • 参考分割