QuadisPD
AI in Healthcare & Digital Biomarkers
A digital biomarker platform for Parkinson's disease that quantifies motor and speech symptoms from video and audio, designed to support clinical assessment and remote monitoring. Currently moving toward clinical pilot testing.
Technologies: Computer Vision, Audio Analysis, PyTorch, Digital Biomarkers, Health-Tech
- Video-based analysis of motor symptoms
- Audio and speech-based symptom signals
- Objective, repeatable symptom scoring
- Built toward a clinical pilot deployment
Virtual Personal Assistant
Multi-Agent AI & Cognitive Systems
A multi-agent cognitive assistant built on a dual-memory architecture, combining short- and long-term memory with spatial reasoning to plan and act across complex tasks.
Technologies: Multi-Agent Systems, LLMs, Dual-Memory Architecture, Spatial Reasoning, Python
- Dual-memory design (working + long-term)
- Spatial reasoning over structured context
- Multi-agent task orchestration
Deep Oscillatory Neural Networks (DONN)
Physics-Grounded Deep Learning & Signal Processing
Physics-grounded autoencoders that use Hopf oscillator dynamics for audio compression and Carnatic music generation, grounding representation learning in dynamical-systems theory.
Technologies: PyTorch, Dynamical Systems, Hopf Oscillators, Autoencoders, Audio DSP
- Hopf-oscillator-based network layers
- Learned audio compression
- Carnatic music generation
- Physics-grounded representation learning
Multi-View 3D Reconstruction + Super-Resolution
3D Reconstruction & Computer Vision
An end-to-end pipeline for multi-view 3D reconstruction with super-resolution, built for a Kaggle competition — pairing feed-forward reconstruction with detail enhancement and Gaussian-splat rendering.
Technologies: VGGT, Real-ESRGAN, nerfstudio (splatfacto), 3D Gaussian Splatting, PyTorch
- Multi-view reconstruction with VGGT
- Super-resolution via Real-ESRGAN
- Gaussian-splat rendering with splatfacto
- Runs across MacBook M2 Pro and T4 GPUs
Stereo 3D Pose Estimation
Computer Vision & 3D
A stereo-vision pipeline that recovers 3D pose from synchronized camera views — detecting 2D keypoints on each view of a calibrated stereo pair and triangulating them into metric 3D positions.
Technologies: Computer Vision, Stereo Vision, 3D Pose Estimation, PyTorch
- 2D keypoint detection on each stereo view
- Stereo calibration and cross-view triangulation
- Metric 3D pose reconstruction from a calibrated pair
Is Deep Learning Enough For Handling X-Ray
Deep Learning & Healthcare
A comprehensive analysis of deep learning models on the Chest X-Ray dataset of 14 common diseases, focusing on the impact of data augmentation and hyper-parameter tuning to improve diagnostic accuracy.
Technologies: Deep Learning, PyTorch, VGG19, DarkNet19, AlexNet, CLAHE, Computer Vision
- Explored the impact of Data Augmentation and Hyper-Parameter Tuning.
- Utilized the Chest X-Ray dataset of 14 Common Thorax Disease Categories.
- Evaluated the performance of various models like VGG19, DarkNet19, and AlexNet.
- Implemented Data Augmentation techniques like Contrast Limited Adaptive Histogram Equalization (CLAHE).
- Tuned Hyper-Parameters such as Learning Rate and Batch Size to improve performance.
An Efficient LLM For Indian Judicial System
AI in Law & NLP
Development of 'LawBot', a novel Large Language Model tailored for the Indian judicial system, designed to overcome the limitations of existing legal tech solutions through custom data scraping and efficient fine-tuning.
Technologies: LLM, Mistral 7B, LoRA, Data Scraping, Python, Fine-Tuning
- Developed a novel 'LawBot' to address limitations of existing legal tech solutions.
- Scraped data from the internet for domain augmentation and classification tasks.
- Employed the Mistral 7B Instruct v0.2, quantised with LoRA adaptors for efficiency.
- Fine-tuned the model sequentially on pre-processed data to enhance legal proceedings.
Generating Music From Text
Generative AI & Music
An ambitious project to generate high-fidelity music directly from text prompts, leveraging the MusicLM architecture and a diverse range of audio datasets for training advanced models.
Technologies: MusicLM, MuLan, Soundstream, word2vec BERT, Generative AI
- Embarking on a project to create music directly from text input.
- Leveraging datasets like AudioSet, AudioCaps, and MusicCaps for training.
- Utilizing advanced models such as MuLan, Soundstream, and word2vec BERT.
- Trained the model and got results that were comparable to the original paper.
Interactive Medical Image Captioning System
AI in Healthcare & Multimodal AI
A system for generating descriptive captions and answering questions about medical images to expedite patient diagnosis and treatment, utilizing the BLIP model on diverse medical datasets.
Technologies: Image Captioning, VQA, BLIP Model, PyTorch, AI in Healthcare
- Developing a system for medical image captioning and question answering.
- Integrating datasets such as MediCat, Path VQA, MIMIC-CXR, and Open-I.
- Incorporating the BLIP Model for accurate captioning and responsive question answering.
- Utilizing advanced multimodal AI techniques such as LLaVa - Vision Language Model.
- Got SOTA results in certain benchmarks
Dawn Of Innovatia: An Animated Short
3D Animation
A captivating 1-minute animated film created in Blender that illustrates the transformative journey of an entrepreneur and highlights the impact of innovation on global progress.
Technologies: Blender, 3D Animation, Storytelling, Video Production
- Produced a captivating 1-minute animated film using Blender.
- Illustrated the transformative journey of an entrepreneur.
- Highlighted the impact of innovation on global progress.
ESummit-Kairos: Dynamic Motion Posters
Motion Graphics & Event Branding
A series of dynamic motion posters designed to encapsulate the essence and theme of the ESummit-Kairos event, contributing to its branding and promotional campaign.
Technologies: Motion Graphics, After Effects, Event Branding, Design
- Engaged in crafting the theme reveal for the ESummit-Kairos event.
- Creating a series of motion posters to encapsulate the essence of the event.
Live and Interactive Portfolio Website
Web Development
A live and responsive personal portfolio engineered with Next.js, featuring dynamic elements from AceternityUI and smooth animations from Framer Motion to showcase my professional journey.
Technologies: Next.js, React, TypeScript, AceternityUI, Framer Motion
- Engineered a live and responsive CV using NextJS and JavaScript.
- Incorporated various dynamic elements using AceternityUI.
- Enhanced user experience with smooth animations using Framer Motion.