Qwen3.8-27B-Object-Detection is a high-capacity vision-language grounding, object detection, and spatial path-mapping workspace powered by the Qwen/Qwen3.8-27B model. The pipeline utilizes native Flash Attention 2 kernels (kernels-community/flash-attn2@v3) to accelerate token decoding and high-resolution visual processing across three dedicated tasks: Object Detection (2D bounding box coordinate prediction), Point Localization (2D keypoint localization), and Spatial Guidance (ordered waypoint and trajectory mapping).
The application is served via a standalone single-page web app (SPA) powered by a FastAPI backend server (gradio.Server) and an interactive dark-mode interface featuring dynamic JSON output parsing, visual bounding-box and waypoint overlays using supervision and PIL, customizable annotation parameters, and history tracking.
-
27B Vision-Language Grounding: Leverages the large-scale
Qwen/Qwen3.8-27Bmultimodal architecture for fine-grained open-vocabulary detection and spatial reasoning. -
Three Dedicated Vision Workflows:
-
Detect: Identifies specific targets and extracts 2D bounding boxes (
bbox_2d) normalized to pixel coordinates. -
Point: Pinpoints exact 2D keypoint coordinates (
point_2d) for referenced objects with visual target indicators. -
Spatial: Maps sequential route waypoints, connecting path segments with directional arrows and legend annotations for trajectory analysis.
-
Flash Attention 2 Kernel Acceleration: Uses
kernels-community/flash-attn2@v3for memory efficiency and accelerated decoding on CUDA hardware. -
Dynamic Visual Overlays: Automatically parses structured JSON model outputs and draws multi-layer masks, custom bounding boxes, and waypoint trails directly over the source image.
-
Interactive Studio SPA: A single-page web interface with real-time JSON display, client-side state synchronization, customizable annotation scales (point radius, box thickness, text size), and history filmstrips.
├── examples/
│ ├── 1.jpg
│ ├── 2.jpg
│ ├── 3.jpg
│ └── 4.jpg
├── src/
│ └── qwen3_8_27b_object_detection/
│ └── __init__.py
├── app.py
├── index.html
├── LICENSE.txt
├── pre-requirements.txt
├── pyproject.toml
├── README.md
├── requirements.txt
└── uv.lock
To set up the Qwen3.8-27B-Object-Detection environment locally, configure your system according to the specifications below. A modern CUDA-enabled GPU with sufficient VRAM is required to run the 27B model.
- Python Version: Minimum Python 3.10.13 or above is required; Python 3.12 or 3.14 is recommended.
- PyTorch Version:
torch==2.11.0or above is required for best compatibility. - CUDA Version: CUDA 13.0 is recommended (
--extra-index-url [https://download.pytorch.org/whl/cu130](https://download.pytorch.org/whl/cu130)), matching the environment used on the live Hugging Face demo.
uv is an ultra-fast Python package and project manager written in Rust. It ensures rapid virtual environment setup and exact dependency synchronization based on the uv.lock file.
Step 1 — Install uv
- macOS / Linux:
curl -LsSf https://astral.sh/uv/install.sh | sh - Windows:
powershell -c "irm https://astral.sh/uv/install.ps1 | iex"
Step 2 — Clone the repository
git clone https://github.com/PRITHIVSAKTHIUR/Qwen3.8-27B-Object-Detection.git
cd Qwen3.8-27B-Object-DetectionStep 3 — Initialize the project and install dependencies
uv syncStep 4 — Run the script
uv run app.py1. Update Package Manager Upgrade your local package manager:
pip install "pip>=26.1.2"2. Install Core Dependencies
Install the deep learning stack, transformer libraries, and vision utilities listed in requirements.txt:
pip install -r requirements.txtaccelerate==1.14.0
peft==0.20.0
transformers-stream-generator==0.0.5
transformers==5.15.0
qwen-vl-utils==0.0.14
sentencepiece==0.2.2
opencv-python==5.0.0.93
torchvision==0.26.0
supervision==0.30.0
matplotlib==3.10.9
einops==0.8.2
spaces==0.51.1
pillow==12.3.0
kernels==0.16.0
gradio==6.22.0
torch==2.11.0
timm==1.0.28
av==17.1.0
Once the server initializes, open your browser to the local address output in your terminal (typically http://127.0.0.1:7860/).
- Upload Input Image: Drag and drop an image into the main canvas workspace, paste from clipboard (⌘/Ctrl+V), or click the upload button in the left rail.
- Select Task Type: Choose your desired operation from the Task Type dropdown:
- Object Detection (Bounding Boxes): Detects object regions and draws yellow boxes with category labels.
- Point Localization (Keypoints): Identifies specific points of interest with circular indicators.
- Spatial Guidance (Path Mapping): Generates ordered waypoint trajectories connected by directional arrows.
- Enter Instruction: Type your prompt (e.g., "Detect all the cars" or "Map a path from the door to the lamp"), or select a prompt from the Quick Prompts chips.
- Run Inference: Click Run Inference or press ⌘/Ctrl + Enter. The annotated image will appear in the main canvas, and the structured JSON response will populate in the output panel.
- License: Apache License 2.0
- GitHub Repository: https://github.com/PRITHIVSAKTHIUR/Qwen3.8-27B-Object-Detection
- Hugging Face Live Space: https://huggingface.co/spaces/prithivMLmods/Qwen3.8-27B-Object-Detection