1 of 12

Neuro-Symbolic Visual Reasoning 🧠

Are we really teaching our children's spatial and visual reasoning?��

"How many red balls are next to the cube but not behind the slide?"

By Aditya & Nikhil

2024204012, 2024201067

Guided by:

Ravi Kiran Sarvadevabhatla,

Makarand Tapaswi,�Mohd Hozaifa Khan

PS: We are asking for sure.

2 of 12

Ishan, a seventh-grader with a knack for storytelling and creative writing, thrived in subjects like English and history. But geometry felt like an alien world. While his classmates breezed through lessons, Ishan stumbled over tasks like:

  • Mental Rotation: Struggling to recognize how a folded net transformed into a 3D shape.
  • Spatial Scaling: Misjudging real-world distances from map scales (e.g., thinking 1:100 meant “steps,” not meters).
  • Perspective-Taking: Unable to predict how a building would look from above or the side.

Story of Ishan

Traditional worksheets and textbook diagrams left him lost. “It’s like everyone else sees a puzzle, and I just see scribbles,” he confided in his teacher, Ms. Patel. His confidence plummeted, and he dodged STEM clubs—a pattern seen in 34% of students with underdeveloped spatial reasoning.

3 of 12

The Problem: Complex Visual Queries

Challenging Queries

  1. "How many red cubes are to the left of the green sphere?”
  1. “Is there a weapon in the scene?”

Current Limitations

Systems fail on nested queries (e.g., “Find a red cube smaller than the sphere to the right of the cylinder”).

Symbolic engines break with noisy inputs

Needed Skills

Understand objects, attributes, and relationships.

For students like Ishan, spatial reasoning isn’t about talent—it’s about tools. Spatial Quest bridged the gap between frustration and mastery, proving that with the right support, every student can decode the world’s hidden geometry.

4 of 12

Our Solution: Spatial quest - A Neuro-Symbolic Visual Reasoning

Neural Perception

Detects objects and recognizes attributes.

Symbolic Reasoning

Applies logical inference and rules.

CLEVR Dataset

Benchmark for complex visual queries.

5 of 12

System Workflow: Detect → Reason → Answer 🧩

1

Input

Visual scene and user query.

2

Detect

Neural network identifies objects and attributes.

3

Reason

Symbolic engine applies logic rules.

4

Answer

System delivers the final result.

6 of 12

Neural Object Detection & Attribute Recognition 🧠

🚂

Training

On sub-sampled CLEVR dataset with 98.7% detection accuracy.

🎁

Objects

Cubes, spheres, cylinders detected precisely.

📲

Attributes

Color, size, and material recognized effectively.

7 of 12

Symbolic Reasoning:

Spatial Rules

Encoded relations: left, right, above, below.

Compositional Queries

Handles multiple conditions logically.

Example Rule

"If A left of B and B left of C, then A left of C."�Nested conditions: "Find objects that are (red AND metallic) OR (small BUT NOT spherical)"

Logical Rules

Propositional logic (AND/OR/NOT)

First-order predicates (∀, ∃ quantifiers)

Domain-specific constraints like physics

8 of 12

Metrics: Accuracy & Speed

Accuracy

Achieved 95.2% on complex CLEVR queries.

Speed

3x faster than end-to-end neural models.

Reasoning Time

Average of 0.7 seconds per query.

9 of 12

Ablations: Neural vs. Hybrid ⚙️

Neural-Only

78% accuracy, struggles with complex queries.

Symbolic-Only

Depends on perfect object detection.

Neuro-Symbolic

95.2% accuracy, robust to noise.

10 of 12

Demo: Interactive Web UI 💻

Explore Visual Reasoning

Children create scenes and test queries interactively.

Drag-and-Drop

Easy scene building with intuitive controls.

Explainability

Visual explanations of the reasoning process.

11 of 12

Conclusion & Next Steps 🚀

Powerful Tool

Enhances visual reasoning skills in kids.

Future Work

Expand to real-world images beyond CLEVR.

Integration

Embed into educational games and learning platforms.

Positive Impact

Supports cognitive development and interactive learning.

12 of 12

Here it all started

Thanks to all 🙏

Nikhil singh AKA 5* Dev

Aditya AKA ProdMan