Skip to content

About

Qwen3.8-27B-Object-Detection is a high-capacity vision-language grounding, object detection, and spatial path-mapping workspace powered by the Qwen/Qwen3.8-27B model. The pipeline utilizes native Flash Attention 2 kernels (kernels-community/flash-attn2@v3) to accelerate token decoding and high-resolution visual processing dedicated tasks.

Topics

Resources

Stars

16 stars

Watchers

0 watching

Forks

Repository files navigation

Qwen3.8-27B-Object-Detection is a high-capacity vision-language grounding, object detection, and spatial path-mapping workspace powered by the Qwen/Qwen3.8-27B model. The pipeline utilizes native Flash Attention 2 kernels (kernels-community/flash-attn2@v3) to accelerate token decoding and high-resolution visual processing across three dedicated tasks: Object Detection (2D bounding box coordinate prediction), Point Localization (2D keypoint localization), and Spatial Guidance (ordered waypoint and trajectory mapping).

The application is served via a standalone single-page web app (SPA) powered by a FastAPI backend server (gradio.Server) and an interactive dark-mode interface featuring dynamic JSON output parsing, visual bounding-box and waypoint overlays using supervision and PIL, customizable annotation parameters, and history tracking.

task qwen_object_detection (1)

Key Features

  • 27B Vision-Language Grounding: Leverages the large-scale Qwen/Qwen3.8-27B multimodal architecture for fine-grained open-vocabulary detection and spatial reasoning.

  • Three Dedicated Vision Workflows:

  • Detect: Identifies specific targets and extracts 2D bounding boxes (bbox_2d) normalized to pixel coordinates.

  • Point: Pinpoints exact 2D keypoint coordinates (point_2d) for referenced objects with visual target indicators.

  • Spatial: Maps sequential route waypoints, connecting path segments with directional arrows and legend annotations for trajectory analysis.

  • Flash Attention 2 Kernel Acceleration: Uses kernels-community/flash-attn2@v3 for memory efficiency and accelerated decoding on CUDA hardware.

  • Dynamic Visual Overlays: Automatically parses structured JSON model outputs and draws multi-layer masks, custom bounding boxes, and waypoint trails directly over the source image.

  • Interactive Studio SPA: A single-page web interface with real-time JSON display, client-side state synchronization, customizable annotation scales (point radius, box thickness, text size), and history filmstrips.

Repository Structure

├── examples/
│   ├── 1.jpg
│   ├── 2.jpg
│   ├── 3.jpg
│   └── 4.jpg
├── src/
│   └── qwen3_8_27b_object_detection/
│       └── __init__.py
├── app.py
├── index.html
├── LICENSE.txt
├── pre-requirements.txt
├── pyproject.toml
├── README.md
├── requirements.txt
└── uv.lock

Installation and Requirements

To set up the Qwen3.8-27B-Object-Detection environment locally, configure your system according to the specifications below. A modern CUDA-enabled GPU with sufficient VRAM is required to run the 27B model.

  • Python Version: Minimum Python 3.10.13 or above is required; Python 3.12 or 3.14 is recommended.
  • PyTorch Version: torch==2.11.0 or above is required for best compatibility.
  • CUDA Version: CUDA 13.0 is recommended (--extra-index-url [https://download.pytorch.org/whl/cu130](https://download.pytorch.org/whl/cu130)), matching the environment used on the live Hugging Face demo.

Running with uv (Recommended)

uv is an ultra-fast Python package and project manager written in Rust. It ensures rapid virtual environment setup and exact dependency synchronization based on the uv.lock file.

Step 1 — Install uv

  • macOS / Linux: curl -LsSf https://astral.sh/uv/install.sh | sh
  • Windows: powershell -c "irm https://astral.sh/uv/install.ps1 | iex"

Step 2 — Clone the repository

git clone https://github.com/PRITHIVSAKTHIUR/Qwen3.8-27B-Object-Detection.git
cd Qwen3.8-27B-Object-Detection

Step 3 — Initialize the project and install dependencies

uv sync

Step 4 — Run the script

uv run app.py

Standard PIP Implementation

1. Update Package Manager Upgrade your local package manager:

pip install "pip>=26.1.2"

2. Install Core Dependencies Install the deep learning stack, transformer libraries, and vision utilities listed in requirements.txt:

pip install -r requirements.txt

Core Requirements List (requirements.txt)

accelerate==1.14.0
peft==0.20.0
transformers-stream-generator==0.0.5
transformers==5.15.0
qwen-vl-utils==0.0.14
sentencepiece==0.2.2
opencv-python==5.0.0.93
torchvision==0.26.0
supervision==0.30.0
matplotlib==3.10.9
einops==0.8.2
spaces==0.51.1
pillow==12.3.0
kernels==0.16.0
gradio==6.22.0
torch==2.11.0
timm==1.0.28
av==17.1.0

Usage

Once the server initializes, open your browser to the local address output in your terminal (typically http://127.0.0.1:7860/).

  1. Upload Input Image: Drag and drop an image into the main canvas workspace, paste from clipboard (⌘/Ctrl+V), or click the upload button in the left rail.
  2. Select Task Type: Choose your desired operation from the Task Type dropdown:
  • Object Detection (Bounding Boxes): Detects object regions and draws yellow boxes with category labels.
  • Point Localization (Keypoints): Identifies specific points of interest with circular indicators.
  • Spatial Guidance (Path Mapping): Generates ordered waypoint trajectories connected by directional arrows.
  1. Enter Instruction: Type your prompt (e.g., "Detect all the cars" or "Map a path from the door to the lamp"), or select a prompt from the Quick Prompts chips.
  2. Run Inference: Click Run Inference or press ⌘/Ctrl + Enter. The annotated image will appear in the main canvas, and the structured JSON response will populate in the output panel.

License and Source

About

Qwen3.8-27B-Object-Detection is a high-capacity vision-language grounding, object detection, and spatial path-mapping workspace powered by the Qwen/Qwen3.8-27B model. The pipeline utilizes native Flash Attention 2 kernels (kernels-community/flash-attn2@v3) to accelerate token decoding and high-resolution visual processing dedicated tasks.

Topics

Resources

Stars

16 stars

Watchers

0 watching

Forks

Contributors

Languages