Theme Icon
Case study · Research Project

IBFET — Large-Scale Code Clone Detection

ResearchSoftware Security
Problem
Comparing billions of lines of source code pairwise for clones exhibits quadratic time complexity O(N^2).
Built
Designed an Index-Based Feature Extraction Technique (IBFET) using MapReduce on Hadoop to build an inverted token index.
Outcome
Indexed 324B+ lines of code, achieving near-linear scalability for massive code-clone detection.

Overview

An index-based feature extraction technique evaluated over more than 324 billion lines of code in a Hadoop distributed environment.

Want something like this built?

I take on serious projects through Megicode — from scope to shipped product. Or browse the rest of the proof first.