Build AI System for (Document or Audio AI) – Offline, Docker, Full Pipeline need Mobile App Development

Contact person: Build AI System for (Document or Audio AI) – Offline, Docker, Full Pipeline

Phone:Show

Email:Show

Location: Nagpur, India

Budget: Recommended by industry experts

Time to start: As soon as possible

Project description:
"I need an experienced AI/ML developer or small team to build an offline-ready AI system for one of the following problem statements (please mention your expertise while applying):

OPTION A: PS-05 – Intelligent Multilingual Document Understanding
You must build a pipeline that:

Parses scanned documents (jpeg/png/pdf/docx) in multiple languages (English, Hindi, Urdu, Arabic, Nepali, Persian)

Detects layout: text, titles, tables, charts, maps, and images

Extracts text with bounding boxes and language

Converts charts/tables/maps/images to natural language summaries

Outputs in proper JSON format

OPTION B: PS-06 – Language-Agnostic Audio Speaker Diarization & Translation
You must build a system that:

Segments speaker turns (Speaker Diarization)

Identifies speakers if enrolled (Speaker ID)

Detects spoken language per segment (Language ID)

Transcribes (ASR) speech to text

Translates text into English

Works offline and outputs CSV/TRN/TXT as per challenge guidelines

Requirements:
Full solution must run offline (no APIs)

Must be packaged in Docker or virtual environment

Models should be open-source or trainable (Whisper, pyannote, LayoutLM, etc.)

Must match file naming, output format, and evaluation structure provided by AI Challenge

Final output must be tested and ready for IIT Delhi demo (if shortlisted)

Deliverables:
Complete codebase with documentation

Docker/Virtual environment ready for offline execution

Output generator in required formats (CSV, TRN, TXT, JSON)

Assistance during evaluation & testing (if shortlisted)

Skills Required:
Python, PyTorch/TensorFlow

ASR/Audio AI (for PS-06)

OCR/Layout detection (for PS-05)

NLP and Translation models

Docker, Linux

Git, JSON, REST APIs (for local inference)

Optional: FastText, Whisper, TrOCR, LayoutLM, HuggingFace

Budget:
Please mention your quote and timeline. Budget can be broken into milestones:

Prototype

Testing

Final Packaging

Bonus if you:

Have worked on Kaggle, AI4Bharat, NeMo, or document/audio pipelines

Can show demos of similar work

To Apply:

Mention which PS (05 or 06) you're confident with

Share your GitHub or project samples (if any)

Suggest any pretrained models you'd recommend using

Propose your timeline

Fast communication and solution-focused mindset preferred." (client-provided description)


Matched companies (0)