
Closed
Posted
I’m expanding a high-performance computing initiative and need seasoned C/C++ talent to squeeze every last cycle out of our code. The core of the work revolves around designing, profiling, and hardening parallel processing routines that will run on Linux/Unix clusters. You’ll be free to choose the right mix of pthreads, OpenMP, MPI, CUDA, or similar frameworks as long as the end result is fast, deterministic, and fully verifiable. Here’s the environment you’ll step into: a mature codebase written almost entirely in modern C and C++, backed by rigorous data-structure design, multithreaded execution, and an SDLC that values deep debugging over quick patches. I handle planning, testing resources, and CI; you focus on writing clean, well-documented modules and driving their performance to the limit. Deliverables • Optimised C/C++ source files for the assigned computation kernels • Profiling reports that demonstrate measurable speed-ups over the baseline • Unit and integration tests that prove functional correctness across threads Acceptance criteria Performance gains must be reproducible on our reference hardware, pass Valgrind and ThreadSanitizer checks, and integrate without regressions into our Git workflow. The engagement is remote, billed hourly, and ideal for developers with 4-12 years of real-world experience who thrive on algorithmic challenges and large-scale optimisation. If that sounds like you, let’s talk.
Project ID: 40652828
22 proposals
Remote project
Active 5 days ago
Set your budget and timeframe
Get paid for your work
Outline your proposal
It's free to sign up and bid on jobs
22 freelancers are bidding on average ₹911 INR/hour for this job

Stepping into your HPC initiative as a seasoned C/C++ developer, my profile extends far beyond coding lines and into the realm of genuinely solving complex bottlenecks and implementing scalable systems. In my vast 27-year career, I've confronted hardware similar to your Linux/Unix clusters, drawing immense value from architectural profiling and parallel processing methodologies. This has streamlined and optimized processes across my projects- whether it's developing Fiber Glass submarines or Drones with Anti-Jamming & RF Communication capabilities as showcased in my rich diverse portfolio. Adept at pthreads, OpenMP, MPI, CUDA, and other such frameworks, I am a firm believer in rigorous testing. Just like your SDLC values deep debugging over quick patches, my work revolves around delivering more than 'just code' but comprehensive performance. I guarantee you measurable speed-ups over the baseline, intricately documented sources for computation kernels and a seamless Git workflow integration without regressions- all verifiable through Valgrind and ThreadSanitizer checks. And let's not forget that I’m experienced enough to guide young developers while staying abreast with the latest technology. Having worked on Zynq-7020, PYNQ FPGA & RF and various Embedded SoCs like STM32 (H7/F4) including Zephyr RTOS and FreeRTOS deterministic tas
₹1,000 INR in 40 days
7.2
7.2

With over a decade's worth of experience in software engineering and coding in C++, I'm ready to embrace your HPC project and optimize it to new heights. With a particular expertise in multi-threaded programming and parallel development, your C/C++ kernel will receive the attention it deserves; focusing on profiling for performance gain and hardening routines for reproducibility. My proficiency in working with frameworks including pthreads, OpenMP, MPI, and CUDA would be value-adding in this endeavor. When it comes to debugging and ensuring robustness, my track record soaring through deep debugging processes rather than opting for quick fixes aligns nicely with your SDLC methodology. Just as your project demands rigorous data structure design, I have refined my skills over time to prioritize clean, well-documented modules and driving their performance to the limit. Indulging in large-scale optimization challenges constitutes one of the intriguing scopes of my professional journey, which has spanned more than 4 years; rendering me capable of delivering software that significantly boosts computational efficiency while adhering closely to acceptance criteria like passing Valgrind and ThreadSanitizer checks along with Git integration without regressions. Shall we begin this stimulating venture together?
₹1,000 INR in 40 days
6.1
6.1

You need a Spring Boot mentor who can guide your architecture decisions instead of only fixing isolated bugs. I help developers build clean backend structures with proper layering, REST API design, database patterns, and maintainable code practices. I have worked with Java Spring Boot projects where service/repository separation, dependency injection, and scalable backend design were key parts of the development process. I can review your package structure, controller design, service logic, repository layer, and suggest improvements based on SOLID principles and real production patterns. I can also explain why a certain approach is better, review feature branches, and help you build a stronger understanding of Spring Boot rather than just delivering quick fixes. Which part of your current application architecture would you like to review first?
₹800 INR in 36 days
4.1
4.1

Hi, I can support your HPC C/C++ parallel development work by optimizing computation kernels, profiling bottlenecks, and improving performance on Linux/Unix cluster environments. My approach will be to first review the current codebase, baseline performance, reference hardware, threading model, data structures, and CI/test workflow. Then I’ll profile hotspots, choose the right parallel strategy, and implement clean, verifiable optimizations using OpenMP, MPI, pthreads, CUDA, or modern C++ concurrency as appropriate. I’m comfortable with C, C++, Linux, parallel processing, multithreading, CUDA, MPI/OpenMP, debugging, profiling, Valgrind, ThreadSanitizer, unit testing, and Git-based development. Deliverables: * Optimized C/C++ source files * Parallel kernel improvements * Profiling report with speed-up evidence * Unit and integration tests * Valgrind/ThreadSanitizer-safe fixes * Clear code comments * Git-ready commits I’ll focus on reproducible performance gains, deterministic behavior, clean integration, and no regressions in your existing workflow. Best regards Ankit
₹750 INR in 40 days
3.8
3.8

As an accomplished Full Stack Developer, my skills and experience extend beyond those mentioned in the project description. Though my profile may differ from your initial expectations, I firmly believe I have the expertise to deliver well-defined, high-performance C/C++ parallel solutions tailored to your needs. Over my 5+ years in the industry, I've consistently helped businesses improve their existing products & scale operations - a skill set that resonates with your project goals. My proficiency in Linux environments combined with my mastery of C programming and software engineering will allow me to tackle this project head-on. In addition, having built complex AI-powered solutions and large-scale systems, I'm well-attuned to algorithmic challenges and adept at squeezing every last cycle out of code. Moreover, my approach aligns with your focus on clean, well-documented code that isn't just functional but also performant. I bring this same level of dedication to parallel processing routines, ensuring not only speed-ups but also reproducibility, verifiability and end-to-end integration via tools like Valgrind and ThreadSanitizer. Above all, I prioritize communication & long-term commitment for successful project delivery – traits that make me an ideal fit for this important project.
₹1,000 INR in 40 days
3.5
3.5

i have working exp with cuda, openCL using c, c++, python, rust. i mostly work on system programming and multicore
₹750 INR in 40 days
2.9
2.9

Thread oversubscription on your Linux clusters can hide latency spikes that break deterministic timing. I'll start by running perf and VTune to map hot spots, then bind OpenMP threads and MPI ranks to separate cores. From there I'll add CUDA kernels where data parallelism pays off, keeping all code under Valgrind and ThreadSanitizer checks. A common mistake is mixing pthread locks with OpenMP schedules, which can cause deadlocks that only appear under heavy load. I always isolate lock scopes and verify them with ThreadSanitizer before merging branches, so regressions never slip through. You’ll receive profiling reports that show clear speed‑up numbers and a test suite that proves each kernel runs correctly across threads.
₹800 INR in 40 days
2.3
2.3

Hello, I have around 14 years of windows development experience using Visual studio/c++ and worked on complex corporate multi-tier architecture projects involving multithreading and multiprocessing. Worked on SunSolaris flavour of UNIX and Linux as well for some telecom projects. Hope I would be able to understand your requirements and deliver them as well. Let's discuss if you feel I'm the right candidate among the bids received. Thanks, and have great day ahead!!!
₹750 INR in 30 days
1.9
1.9

I'm Geetam, a seasoned software engineer with a focus on performance optimization. Having served as the sole technical partner for numerous projects, I possess the unique ability to manage end-to-end development tasks - an invaluable asset given the depth and complexity of your project. My proficiency in C/C++, Linux, and debugging makes me a natural fit for your parallel programming initiative. In my decade-long career, I have honed my skill set through real-world experience in algorithmic challenges and large-scale optimization projects - facets that are central to the success of your endeavour. I am confident with frameworks such as pthreads, OpenMP, MPI, CUDA, and more, which are imperative for squeezing out every last cycle from your code on Linux/Unix clusters. My work is usually backed by rigorous data-structure design and multithreaded execution which aligns seamlessly with your mature codebase. For deliverables, beyond providing optimized C/C++ source codes and profiling reports that manifest measurable sped-up versions of your baseline code, I am proactive in ensuring functional correctness across threads by formulating thorough unit and integration tests. To reinforce this point, I endeavor to pass Valgrind and ThreadSanitizer checks while integrating without any regressions into the Git workflow.
₹900 INR in 40 days
1.6
1.6

Hello, I’m Bharghav, and I bring 10 years of experience in matching job skills, particularly in C Programming, Linux, and C++ Development. With a solid background in high-performance computing, I'm well-prepared to tackle the challenges presented in your project. I understand that your initiative focuses on optimizing and hardening parallel processing routines for Linux/Unix clusters. My approach will involve diving deep into your mature codebase, utilizing the appropriate frameworks such as pthreads, OpenMP, or MPI to enhance performance while ensuring clean and maintainable code. I will deliver optimized C/C++ source files for the computation kernels, backed by comprehensive profiling reports and rigorous testing to validate functionality across threads. Let’s connect in a chat to discuss your project in detail and explore how I can contribute to its success. Best regards, bhargav922002
₹875 INR in 3 days
1.3
1.3

As a seasoned software engineer with over a decade of experience, I've not only honed my skills in C/C++ but have also delved into crucial parallel programming frameworks like pthreads, OpenMP, MPI, and CUDA - key technologies for your HPC project. My familiarity with your tech stack and the expectations surrounding algorithmic challenges and large-scale optimization make me a strong fit for your team. I understand that performance is at the core of your initiative, and as such, I give it paramount importance within my own work. My proficiency in designing, profiling, and hardening parallel processing routines on Linux/Unix clusters will contribute immensely in understanding and addressing your project's objectives. Additionally, my deep debugging approach aligns well with the SDLC you mentioned, ensuring clean modules that not only drive performance but can be easily maintained. Lastly, my remote working proficiency comes as an added advantage; accustomed to remote setups, I'm comfortable collaborating effectively while delivering my work on time. As we embrace the new normal in remote working conditions, it’s vital to have freelancers who can adapt quickly without compromising on quality or delivery milestones. I assure you of consistent availability and weekly meetings for status updates throughout the project duration. Let's talk further about how we can make your HPC initiative even more successful!
₹750 INR in 40 days
1.4
1.4

When you push parallel loops onto a NUMA node, cache line bouncing can erase any speed gain. I'll pin threads to cores, use first-touch allocation, and run perf to confirm locality improvements. The profiling report will include reproducible timings on your reference hardware, and the code will pass Valgrind and ThreadSanitizer checks. A common pitfall is assuming more threads always mean faster runs; lock contention often flips the curve. I'll write unit tests that fire each kernel under varied thread counts to catch nondeterministic bugs early. Ready to start immediately and push the modules into your Git CI pipeline.
₹1,000 INR in 40 days
0.0
0.0

I can help with the C/C++ parallel optimisation work, especially where the deliverable needs profiling evidence rather than just code changes. My first milestone would be: 1. reproduce the baseline on your reference build, 2. identify hot paths with profiling, 3. optimise the assigned kernel using the right tool for the shape of the workload: OpenMP/MPI/CUDA/pthreads, 4. add correctness tests around threaded behavior, 5. provide a before/after profiling note and any Valgrind/ThreadSanitizer findings. Live proof page showing my working-demo style: https://rufael-live-freelancer-demos.rufaelaman.chatgpt.site#validator If you send the target kernel and baseline timings, I can start with the smallest measurable speed-up and build from there.
₹925 INR in 2 days
0.0
0.0

Your approach of combining a mature codebase with rigorous profiling, debugging, CI, and reproducible benchmarks is an excellent idea. I can take ownership of the optimization work and deliver clean, documented, tested C/C++ modules with demonstrable speed improvements. Could you provide the existing profiling data or baseline benchmark results for the first assignment?
₹1,000 INR in 40 days
0.0
0.0

I think your idea is excellent because the acceptance criteria clearly connect optimization work with measurable and verifiable engineering results. I can implement the required C/C++ optimizations and provide profiling evidence, comprehensive tests, and validation through Valgrind and ThreadSanitizer. Would you like me to begin with a specific bottleneck or perform an initial profiling pass across the relevant codebase?
₹1,000 INR in 40 days
0.0
0.0

We have over 5 years experience with similar projects in high-performance computing. You're looking to optimize parallel processing routines in a mature C/C++ codebase while ensuring performance gains are reproducible and verifiable. I would approach this by first analyzing the existing codebase to identify bottlenecks, then implement the most suitable frameworks like OpenMP or MPI for parallelization. My focus will be on writing clean, well-documented modules that enhance performance while integrating seamlessly into your Git workflow. I will ensure thorough profiling to demonstrate measurable speed-ups and rigorous testing for functional correctness. Deliverables: • Optimized C/C++ source files for the assigned computation kernels • Profiling reports that demonstrate measurable speed-ups over the baseline • Unit and integration tests that prove functional correctness across threads I am happy to share relevant examples of my work. Would you like to discuss any specific areas of the codebase to focus on first? Regards, RyanF172
₹750 INR in 7 days
0.0
0.0

I can help optimize your high-performance computing initiative by enhancing your C/C++ code for maximum efficiency. The focus on rigorous data-structure design and multithreaded execution aligns perfectly with my skill set. I appreciate that you value deep debugging and a mature codebase. This approach ensures that the performance improvements I implement will be measurable and reproducible, meeting your acceptance criteria. With my experience in using frameworks like OpenMP and MPI, I can write clean, well-documented modules that integrate smoothly into your workflow. I’m ready to tackle algorithmic challenges and drive performance to the limit. Let’s chat about your project and the best approach moving forward. At the very least, you’ll get a free consultation. Regards, DeanSaliegh I have done similar work: Modern Real Estate Investment Website
₹750 INR in 7 days
0.0
0.0

I am c/c++ developer with more than 10 years experience so i am confident will make best qualify for this task
₹1,250 INR in 40 days
0.0
0.0

Hi there, Looking at the current bids, it seems you are receiving many proposals from generalist web or app developers. My focus is strictly on low-level system performance, algorithmic challenges, and deep C/C++ optimization. I specialize in writing, refactoring, and benchmarking C and C++ code. To squeeze every last cycle out of the hardware, I frequently drop down to x86_64 Assembly and utilize SIMD/AVX2 vectorization instructions. I work natively in the Arch Linux terminal, meaning the mature, heavily-debugged environment you described is exactly where I am most productive. In my independent engineering work—such as developing the 'Arbitrary' library entirely from scratch and building the 'Scout' command-line framework—I prioritize rigorous data-structure design and deterministic execution. I am fully accustomed to proving functional correctness, thread safety, and memory integrity using Valgrind and ThreadSanitizer before pushing anything to Git. Whether the solution requires threading via pthreads/OpenMP or rewriting core loops for deep vectorization, I can deliver clean, heavily profiled modules that show measurable speed-ups without introducing regressions. I'd love to take a look at a baseline profile. Are the primary bottlenecks in your current computation kernels mostly compute-bound, or are you hitting memory bandwidth limits?
₹1,000 INR in 38 days
0.0
0.0

Hello, I’m a Senior C++ Developer with 4 years of professional experience, primarily working on performance-sensitive applications in the gaming industry. I have strong hands-on experience with modern C++, multithreading, STL, data structures, memory management, debugging, performance optimization, and Git-based development workflows. Your project is a great match for my skills. I can help with: - C/C++ performance optimization and profiling - Multithreading and concurrency - Algorithm and data-structure optimization - Memory and CPU performance improvements - Debugging race conditions and complex runtime issues - Valgrind and ThreadSanitizer analysis - Unit and integration testing - Linux-based C++ development I’m comfortable working with an existing codebase, identifying bottlenecks, implementing clean solutions, and validating measurable performance improvements without compromising correctness. I’m new to this freelancing platform, but I have 4 years of real-world professional C++ development experience. I’m looking to build a long-term client relationship and can start with a small task or module so you can evaluate my skills and work quality. I’d be happy to discuss your requirements and get started. Best regards, Raj
₹1,000 INR in 40 days
0.0
0.0

Nagpur, India
Member since Aug 17, 2026
$25-50 USD / hour
$25-50 USD / hour
$10-30 USD
$500 USD
$1500-3000 USD
$10-30 USD
€8-30 EUR
$10-12 USD
$30-250 USD
₹600-1500 INR
₹12500-37500 INR
₹600-1500 INR
₹1500-12500 INR
₹1500-12500 INR
₹750-1250 INR / hour
₹1250-2500 INR / hour
$15-25 USD / hour
₹120-240 INR / hour
$20-250000 AUD
₹1500-12500 INR