§ 01/ Engineer

Abhinav
Dubey.

AI Engineer.

§ 02/ Experience
TEKsystems Global Services
4+
years
Shipping production systems at TEKsystems Global Services, from trainee to Senior Engineer.
  • Senior Software Engineer, AI
    Apr 2025 to presentCurrent
  • Software Engineer, AI
    Apr 2024 to Mar 20251 yr
  • Associate Software Engineer, AI
    Sep 2022 to Mar 20241 yr 6 mo
  • Technical Trainee
    Feb 2022 to Sep 20227 mo
§ 03/ About

I'm an AI engineer with 4+ years of experience shipping systems for Fortune 500 clients, including Ahold Delhaize, State Farm, and the Automobile Club of Southern California. Before AI, I worked on distributed systems, so I sit at what I think is the right intersection: backend rigor that makes AI systems reliable, and AI craft that makes them useful.

Some of what I've shipped: Claude Code plugins for developer workflows, a ServiceNow ticket resolution agent that closes L1 tickets end to end, multi source RAG pipelines that pull from Jira, Confluence, and internal docs in a single pass, and fine tuned small language models on domain specific support data so they respond in the client's voice.

I'm a Claude Certified Architect from Anthropic's Foundations program. I also hold Google Cloud's Generative AI Leader certification and TEKsystems' 2023 Rising Star award.

Off duty, I play football on weekends with friends and read novels every chance I get. My dream is to travel the world one country at a time, Europe first. Looking forward to it.

Based in
India
Working at
TEKsystems Global Services
Since Feb 2022
Off duty
Football, novels, and a slow plan to travel Europe.
Resume
Download PDF
§ 04/ Writing
The Fastest Way to Speed Up an AI Model? Make It Skip Parts of Its Own Brain
Medium2026, 7 min
The Fastest Way to Speed Up an AI Model? Make It Skip Parts of Its Own Brain

On skip-attention, sliding windows, and why on-device speed comes from measurement, not theory.

Here is a weird idea. If you want an AI model on your phone to respond faster, one of the best things you can do is have it skip steps entirely in some of its layers. On TTFT, sliding windows, skip attention, and letting the device tell you what actually wins.

Read on Medium
Understanding the Self-Attention Mechanism in Transformers (with PyTorch Code)
Medium2025, 4 min
Understanding the Self-Attention Mechanism in Transformers (with PyTorch Code)

Q, K, V from scratch. What the softmax is actually softmax-ing over.

A step by step walk through self-attention with runnable PyTorch. Query, Key, Value; the scaled dot-product; softmax normalization; masking; multi-head. If BERT and GPT feel like magic to you, this is where the magic actually happens.

Read on Medium
§ 05/ Stack
Languages and Backend
PythonPython
GoGoLang
Node.jsNode.js
PyTorchPyTorch
FastAPIFastAPI
ReactReact
REST
GraphQLGraphQL
PostgreSQLSQL
RedisRedis
01 / 05
§ 06/ Selected work
3 projects
TokenStash
Claude Code plugin, personal tool2026
TokenStash

A Claude Code plugin that caches project context locally, so Claude starts every session already knowing your codebase. I got tired of re-explaining my repos, so I built it for my own workflow.

Claude
View on GitHub
Intelligent Contract Analysis Agent
RAG, learning demo2025
Intelligent Contract Analysis Agent

A demo that reads dense legal contracts and flags risky clauses, compliance gaps, and negotiation points. I built it to see how far off-the-shelf RAG and GPT-4 get before they miss the important stuff.

LangChainSupabaseStreamlitPython
View on GitHub
YouTube Assistant
LangChain, afternoon build2024
YouTube Assistant

Paste a YouTube URL, ask questions, get answers grounded in the transcript. An afternoon build to feel out the whole LangChain retriever surface: loaders, splitters, embeddings, FAISS.

PythonStreamlitLangChain
View on GitHub
§ 07/ Recognition
§ 08/ Contact
Email
dubey.abhinav76@gmail.com
LinkedIn
/in/abhinavdubeyad9
GitHub
GitHub
/abhinavdubeyad9
Medium
Medium
@dubey.abhinav76
Location
India