Projects

AI and Tagore Project

The Quiet Side of AI:
Discovering Its Potential Through Tagore

By Sankha Kar
Chief Photographer, School of Mass Communication, KIIT (Deemed University), Odisha, India
Former Deputy Photo Editor, Gulf News, Dubai, UAE
Knight Fellow (2024–25), School of Visual Communication, Ohio University, USA

Subha

When people talk about Generative AI, or simply Artificial Intelligence, the conversation usually begins with either laughter or fear. They point to viral AI images they have seen online: the Pope in a white puffer jacket dribbling a basketball, President Trump being arrested on a busy street, Elon Musk confidently walking the runway at Paris Fashion Week, or Bollywood actors reimagined as superheroes and mythological figures. These images spread across social media within seconds and entertain millions of people, but they also create confusion. Many viewers cannot easily distinguish between what is real and what is artificially constructed. For a large number of people, these viral images become the definition of AI, a tool for jokes, exaggeration and digital chaos. Behind the humour lies a deeper anxiety that is difficult to ignore. People worry about deepfakes, manipulated videos, fabricated news, and the general erosion of trust in the truthfulness of images.

As someone who has worked in photography for more than thirty-five years, I have felt this impact very directly. Generative AI has entered the visual world so quickly and with such intensity that at times it feels as though the very foundation on which photorealism stands has begun to shake. The fear is not only about the arrival of fake pictures, but also about the effect on an entire profession that has always depended on trust, authenticity and the human presence behind the camera.

The Noise Around AI

AI entered our lives with too much noise, too quickly. Every platform is filled with strange, humorous or surreal images, and suddenly everyone has an opinion. Many describe AI as dangerous, dishonest or shallow. It has become very easy to blame AI for the loss of authenticity. But throughout my career, I have seen that image manipulation is far older than this new technology.

Long before AI existed, photographers in the nineteenth century were staging and combining negatives to create scenes that never truly happened. A well-known example is the Cottingley Fairies photographs from 1917, in which two young girls fooled the world, even convincing Sir Arthur Conan Doyle, by using simple cut-out paper figures. In the 1930s and 1940s, the Soviet government under Stalin often erased political enemies from official photographs, literally rewriting history with scissors, ink and deliberate omission. Even war photography is not free from controversy. During the Crimean War, the photographer Roger Fenton was accused of moving cannonballs to create a more dramatic battlefield scene in his photograph titled The Valley of the Shadow of Death. These examples make it clear that photography has never been completely free from manipulation.

In the digital era, the conversation did not become easier. Newsrooms around the world argued about cloned pixels, misleading composites and staged scenes. Editors debated what qualified as ethical practice long before AI entered the picture. Like nuclear technology, Generative AI is simply potential. It can illuminate or destroy depending on who uses it, and the responsibility always rests with the human being in control. This understanding stayed with me. The problem has rarely been the technology itself. It is the intention of the person who uses it. AI is only the newest tool in a long line of tools that can either reveal truth or distort it.

Turning to Tagore

In the middle of the noise surrounding AI, I felt the need to turn inward. Rather than chasing trends or reacting to the latest AI spectacle online, I returned to something familiar and grounding: literature. More specifically, I returned to the works of Rabindranath Tagore. Tagore’s writings have shaped my imagination since childhood. He helped define modern Bengali identity through poetry, music, short stories and a profound understanding of human emotions.

Among his many works, the short story Subha has always held a special place in my heart. Written in 1892, Subha tells the story of a mute girl whose entire emotional life is expressed through her eyes, gestures and silence. It is a story of loneliness, quiet struggle and unspoken communication. These delicate themes felt perfectly suited for a different kind of experiment with Generative AI. I wanted to see whether AI, often associated with loud and exaggerated imagery, could engage with a story that breathes softly, almost in whispers.

The challenge extended beyond emotional subtlety. Bengali life in the late nineteenth century is full of cultural details that AI often misrepresents. Wedding rituals such as the topor worn by Bengali grooms, the specific atpoure draping style of the Bengali sari, the texture of handwoven cotton fabrics, the layout of rural homes and the gentle, faded colours of everyday life are rarely captured accurately in AI-generated images. Because these elements are underrepresented in many image datasets, AI tends to default to generic or modern visuals that distort the cultural truth of the story.

This is why Subha felt like the right choice. It offered both a challenge and a responsibility. The emotional depth of the narrative required sensitivity, while its cultural richness demanded accuracy. I was not trying to replace photography but to understand whether this new tool could support imagination in the same way that new techniques and technologies have influenced art throughout history. My goal was to see whether AI could help rebuild a world that no longer exists while still honouring its emotional truth.

Planning and Pre-Production

Before creating any visuals, I began by reading Subha again and again, identifying its key themes and the cultural elements that shaped the story's world. Understanding the rhythm of the narrative helped me break it into smaller moments, such as turning points, silent expressions and emotional beats that could be translated into visual frames. I created a rough storyboard and mood board to maintain consistency. This step was important because consistency is one of AI’s greatest weaknesses, and without clear planning the visual flow becomes unstable.

Character identification was the next step. Subha’s appearance, body language, facial expression and general demeanour needed to remain stable throughout the project. I created reference sheets to prevent visual drift. I also researched locations central to the story: rural homes, riverbanks, courtyards, temple paths and wedding spaces. These details formed the structure of the environments I wanted to generate, ensuring they felt true to the time and culture in which Subha lived.

Research

Research became the backbone of this project. Subha is rooted in a culture and a historical moment that cannot be recreated from memory alone. A quick search on Adobe Stock or similar platforms for terms like Bengali wedding or nineteenth century Bengal girl reveals how poorly these visuals are represented. The problem is not always with AI itself but with the lack of accurate source material. Without proper research, any AI-generated image risks becoming a vague or incorrect representation.

I explored picture archives, articles and books on Bengali attire, rituals and village life. I revisited the films of Satyajit Ray, whose detailed and culturally accurate visual style became an essential reference. I studied old photographs, paintings and illustrations from museums and the National Archives of India. These materials helped me understand the textures of fabric, the structure of jewellery, the layout of rural homes and the emotional environments in which people lived. Academic papers supported this by helping me understand the social dynamics and gender expectations of nineteenth century Bengal. This layered research ensured that every AI frame carried emotional truth and cultural depth.

Creating AI Prompts for Consistency

Prompts became the foundation of the entire visual narrative. The structure, clarity and specificity of each prompt directly influenced the accuracy of the final image. I built a core set of prompts that defined characters, locations, lighting and emotional tone. Each initial output helped me refine these prompts further. Describing characters in detail, including age, ethnicity, clothing and facial features, helped maintain stability across images. Environmental cues rooted the visuals in rural Bengal. Describing lighting conditions and emotional cues brought the images closer to the tone of Tagore’s narrative.

This process taught me that prompt writing is not merely technical. It is a form of direction, similar to guiding a camera crew or briefing actors. You must tell the AI where to stand, what to observe, and how to feel.

Visualising Subha

The challenge of bringing Subha to life began with defining her physical and emotional world. My first attempt used Adobe Firefly, but the images looked distinctly modern. Hairstyles, facial structures and fashion choices felt disconnected from nineteenth century Bengal. This taught me that AI does not understand culture. It only repeats patterns from its training data. If those patterns are inaccurate, the results will be inaccurate.

Visualising Subha



Realising this, I moved to Midjourney, which allowed for deeper control. Through multiple iterations, I created a version of Subha that felt believable. Then I built the other characters: her gentle father, her conflicted mother, and Pratap, the boy who briefly enters her world. I generated variations of each character from different angles to create a cinematic flow. One of the most challenging tasks was portraying Subha across different stages of her life. Maintaining consistency in facial structure, posture and emotional tone required careful attention. Gradually, I created a sequence of images showing her growth from infancy to a young bride. This process was not simply about generating images. It was about recreating an inner world. AI became a tool that reflected my intention, not a replacement for it.

The Challenges and Limitations of Generative AI

Working with Generative AI revealed its weaknesses alongside its strengths. Creating multiple characters in one frame often resulted in merged faces or unusual body proportions, especially in complex scenes. Clothing accuracy was another major issue because AI consistently struggled to represent traditional Bengali clothing. Wedding elements such as the topor, mukut or the bride’s forehead chandan design were repeatedly incorrect or missing, and I had to manually add many details using Generative Fill tools.

AI is highly literal. In one case, a simple descriptive word, small, resulted in a calf being generated at a bizarre miniature scale. Removing that word corrected the entire scene. This taught me that prompt precision is essential.

Inconsistency remains the greatest challenge. AI tools generate unique visuals each time, which makes maintaining continuity difficult. Characters change slightly from one frame to the next. Backgrounds shift unexpectedly. Lighting fluctuates and perspectives do not match. Fine details often fail to stay stable. All of this increases the need for manual editing. The process made it clear that strong artistic direction is still essential for storytelling.

Using Generative AI Effectively

I learned that working with AI is an iterative process. The first prompt rarely produces the right image. Refinement is essential. Reference images help compensate for cultural gaps in AI training data. Structured prompts provide a foundation for consistency. Emotion, lighting and camera perspective influence the final outcome. Too many instructions can confuse the system. Upscaling can sometimes remove natural softness, so manual corrections become necessary. Ultimately, AI works best when paired with patience, precision and cultural awareness.

Best Practices for Prompting

Effective prompting became a form of visual direction. Characters, environments and moods must be defined with clarity. Emotional cues guide expression. Repetition of phrasing helps control variability. Specific descriptions prevent stereotypical or inaccurate visuals. Prompting becomes a collaboration between the human imagination and the machine’s interpretation. The success of AI storytelling depends entirely on the clarity of this conversation.

What This Journey Taught Me

Looking back, I see that this project was not simply an experiment with technology. It was a journey of patience, cultural responsibility and artistic discovery. Reimagining Subha taught me that AI is not the storyteller. The human remains in control, shaping each decision with intention and care.

The most striking realisation is how rapidly AI is evolving. I completed this project in May 2025. In just seven months, the technology has changed dramatically. Tools that once struggled with cultural accuracy are improving. Features that once felt experimental are now becoming ordinary. This rate of change forces all of us in photography, visual storytelling and education to rethink what is possible.

Rather than resisting this shift, I have chosen to grow with it. After completing Subha, I started exploring how AI can support photography education. I am now creating teaching materials in which complex photographic lessons, such as lighting, composition and lens behaviour, are explained through cartoon characters and comic strips generated by AI. These visuals make difficult ideas simple and accessible, especially for beginners. What once required costly equipment or long demonstrations can now be illustrated quickly through AI-generated storytelling.

This project, which began as a personal experiment, has opened a new path for me. It has shown me that AI will not replace photography or diminish its value. Instead, it can expand how we teach, learn and imagine. When I completed the adaptation of Subha, I did not feel victorious or technologically advanced. I felt peaceful. The journey was not about proving what AI can or cannot do. It was about discovering what I can do with AI without losing the essence of who I am as a photographer and a storyteller.