Build AI System for (Document or Audio AI) – Offline, Docker, Full Pipeline need Mobile App Development
Contact person: Build AI System for (Document or Audio AI) – Offline, Docker, Full Pipeline
Phone:Show
Email:Show
Location: Nagpur, India
Budget: Recommended by industry experts
Time to start: As soon as possible
Project description:
"I need an experienced AI/ML developer or small team to build an offline-ready AI system for one of the following problem statements (please mention your expertise while applying):
OPTION A: PS-05 – Intelligent Multilingual Document Understanding
You must build a pipeline that:
Parses scanned documents (jpeg/png/pdf/docx) in multiple languages (English, Hindi, Urdu, Arabic, Nepali, Persian)
Detects layout: text, titles, tables, charts, maps, and images
Extracts text with bounding boxes and language
Converts charts/tables/maps/images to natural language summaries
Outputs in proper JSON format
OPTION B: PS-06 – Language-Agnostic Audio Speaker Diarization & Translation
You must build a system that:
Segments speaker turns (Speaker Diarization)
Identifies speakers if enrolled (Speaker ID)
Detects spoken language per segment (Language ID)
Transcribes (ASR) speech to text
Translates text into English
Works offline and outputs CSV/TRN/TXT as per challenge guidelines
Requirements:
Full solution must run offline (no APIs)
Must be packaged in Docker or virtual environment
Models should be open-source or trainable (Whisper, pyannote, LayoutLM, etc.)
Must match file naming, output format, and evaluation structure provided by AI Challenge
Final output must be tested and ready for IIT Delhi demo (if shortlisted)
Deliverables:
Complete codebase with documentation
Docker/Virtual environment ready for offline execution
Output generator in required formats (CSV, TRN, TXT, JSON)
Assistance during evaluation & testing (if shortlisted)
Skills Required:
Python, PyTorch/TensorFlow
ASR/Audio AI (for PS-06)
OCR/Layout detection (for PS-05)
NLP and Translation models
Docker, Linux
Git, JSON, REST APIs (for local inference)
Optional: FastText, Whisper, TrOCR, LayoutLM, HuggingFace
Budget:
Please mention your quote and timeline. Budget can be broken into milestones:
Prototype
Testing
Final Packaging
Bonus if you:
Have worked on Kaggle, AI4Bharat, NeMo, or document/audio pipelines
Can show demos of similar work
To Apply:
Mention which PS (05 or 06) you're confident with
Share your GitHub or project samples (if any)
Suggest any pretrained models you'd recommend using
Propose your timeline
Fast communication and solution-focused mindset preferred." (client-provided description)