Despite having significantly fewer, 40-billion-parameter model with a context window of 128K tokens, Achieves:
81.4% on SWE-Bench Verified,
49.9% on BigCodeBench,
81.1% on LiveCodeBench v6.
🟢 How it works?
The model uses the "code-flow" technique - learning from the evolution of repositories and commits, and is divided into 2 branches.
Dense Models: Base and Instruct versions for fine-tuning and following instructions
Loop Models: an optimized version with maximum VRAM efficiency (int4 can run on 3090\4090)
LoopCoder architecture uses a cyclic transformer structure, where the same model parameters are used in 2 consecutive data processing passes.
In the first pass, the model processes embeddings through its layers, taking into account the positions of words.
In the second pass, the model simultaneously uses two types of attention: global attention, which refers to all information from the first pass to understand the general context, and local attention, which only looks at the previous words in the second pass to preserve the text sequence.
Both types of attention are combined using a mechanism that decides how much weight to give to the global context and how much to the local sequence.
Modified MIT License, Project page, Technical report, Model set, GitHub
#AI #ML #LLM #IQuest #QuestResearch
••••••••••••••••••••••••••••••••••••••••••••••••••••
🤖 Data Science, ML & Big Data with @DataXplore
