Private GPU Performance Mentorship

Your code. Your bottleneck. One-on-one with a GPU performance engineer.

This is not a course. There is no syllabus, no recorded lectures, no cohort to keep up with. You bring the kernel that won’t go faster, the port that stalled, or the profile you can’t read โ€” and we work on it together, live, until you can do it yourself.

Led by Dr. Mandar Gurav โ€” PhD, IIT Bombay ยท Ex-Intel, NVIDIA, CDAC, IITBombay ยท 16+ years in CUDA, OpenMP, MPI, and GPU performance engineering. The same person who does SVACL’s audit and parallelization work. Not a teaching assistant.

[ Book a diagnostic call โ†’ contact@svacl.com ]


Who this is for

  • The engineer with a stalled GPU port. The CPU version works. The CUDA version is slower, or wrong, or both โ€” and the tutorials all stop right before your problem starts.
  • The researcher whose kernel is slow and doesn’t know why. You’ve read the occupancy calculator. You’ve tried more blocks. The number won’t move.
  • The engineer who’s been handed a GPU codebase. Someone left, you inherited it, and you need to become the person who understands it โ€” quickly.
  • The team lead who needs one person brought up to speed fast. Sending them to a generic 3-day course won’t do it. This will.

Who this is not for

Read this part carefully โ€” it saves us both time.

  • This is not coursework or assignment help, and not thesis writing.
  • This is not a debugging service. I don’t take your repository away and fix it between sessions. (If that’s what you need, you want a GPU Performance Audit โ€” different service, different page.)
  • This is not an introduction to programming. You should be comfortable in C, C++, Fortran, or Python before we start.

If you’re a beginner looking for a structured path from zero, mentorship is the expensive way to get it. Say so up front and I’ll point you somewhere better.


How it works

1. Diagnostic call โ€” โ‚น5,000, 60 minutes. You show me the code, the profile, or the problem. I tell you straight whether mentorship is the right tool, what we’d work on, and how many sessions it will realistically take. If we go ahead, this โ‚น5,000 is credited against your package.

2. A written plan. Before you pay for anything else, you get a short plan: what we’ll cover, in what order, and what “done” looks like.

3. Prepaid sessions โ€” 60 minutes each, on your real code. Screen shared, profiler open, your workload on your hardware. We measure, we locate the bottleneck, we prove the fix moved the number. Same method SVACL uses on paid engagements โ€” you’re just in the driver’s seat.

4. Notes after every session. What we found, what we changed, what to try before next time. Yours to keep.


The method

Most “optimization help” is someone looking at code that seems slow and guessing. We don’t guess.

  • Measure โ€” profile the real workload on real hardware.
  • Locate โ€” find the true bottleneck, backed by roofline and occupancy data. It’s usually not where you think.
  • Prove โ€” fix it, then benchmark before and after, against your reference output.

By the end, the goal isn’t that your kernel is faster. It’s that you know how you made it faster โ€” and can do it again without me.


Pricing

All sessions are 60 minutes. Packages are prepaid and valid for 90 days.

EngagementPrice (India)Price (International)Includes
Diagnostic callโ‚น5,000$100 60 min, credited toward any package below
Working professional (self-funded)โ‚น20,000$500 6 sessions
Student / self-learnerโ‚น12,000$250 4 sessions (2 slots per month)
Academic institution (institute pays)โ‚น5,000/hr$100/hr 8-hour minimum
Company-funded (your employer pays)From โ‚น10,000/hrFrom $200/hr4-hour minimum, or a monthly retainer. Get in touch.

Special offer – upto 40% off for the first 5 mentees. Ends 31 August. Same sessions, same person, introductory price. In exchange I’ll ask you for honest feedback and, if the work was good, a testimonial.

Capacity: I take a maximum of 4 mentorship hours per week. The rest of the week belongs to audit and parallelization clients. When the slots are full, they’re full.


Company-funded engagements

If your employer is paying, this is consulting with a teaching wrapper โ€” and it’s priced that way.

You get the same 1:1 sessions, but on your production code, under NDA, with async code review between sessions. Typical shape: one engineer, four hours a month, working through a real performance problem while learning to do it independently. Invoiced to the company, GST included.

Teams often start here and move to a GPU Performance Audit once it’s clear the bottleneck is bigger than one kernel.

Write to contact@svacl.com with a one-paragraph description of the problem and we’ll take it from there.


Common questions

What do I need before the first session? Code that builds, a workload that runs, and access to the GPU you actually care about. Nsight Systems and Nsight Compute installed, if you can.

Can we do this in Marathi or Hindi? Yes. English, Hindi, or Marathi โ€” whichever you think in.

Do you sign NDAs? Yes, routinely. Standard practice for company-funded engagements.

What if the first session tells me I don’t need mentorship? Then I’ll say so, and you’ll have spent โ‚น5,000 to avoid spending โ‚น20,000. That’s a good outcome.

Can I buy a single hour? No. Packages only. One-off hours turn into a stream of “quick questions” and serve nobody well.


Ready?

Tell me what you’re working on โ€” the slow kernel, the stalled port, the profile that makes no sense. One paragraph is enough.

contact@svacl.com ยท (+91) 9373881607