I wonder if the big AI companies train their public AI models on their internal code or feel protective of their IP and keep it to themselves.

submitted by
6
50

Back to main discussion

Most AI models at this point won’t see significant gains from training on such a small sample of code.

You don’t need a whole corporation’s code to make a functional model, you need the whole world’s.

Adding a tiny bit of your own company’s code to the mix doesn’t really do anything to change the model much, so they generally won’t do it for that reason. Tons of training costs, the only benefit is that the model is very very very slightly fine tuned to kinda sorta produce code that’s maybe possibly a little more stylistically similar to yours.

We’re talking about huge companies with unfathomably huge codebases written by tens of thousands of people. They control significant chunks of the world’s code. It would be stupid not to at least include it in an internal model.

As big as some individual corporations are, the world (including every other massive corporation) is much bigger.




Insert image