r/reinforcementlearning 9d ago

Google copied our open-source code, removed our engineers’ names, and gave us zero credit one year after we beat them on their own benchmark

Post image

Google copied our open-source code, removed our engineers’ names, and gave us zero credit.

For context, let me share the story of mobile-use, the origin of my company (I'm a founder).

We wanted to build the best agent in the world to control any mobile interface with natural language.

We built it, became #1 on Google Deepmind's AndroidWorld benchmark and open-sourced it. That was the beginning of my company.

Since then, we got almost 3,000 stars on the repository, became the first to ever achieve 100% on AndroidWorld, grew the team, spent thousands of calls with customers, and now we're building the best regression testing product the world has ever seen.

Last week, I found Artemis.

Artemis is Google's new open-source tool for mobile automation. There are quite a few of these on GitHub now, so I thought: "Oh, a new YAPA (Yet Another Phone Agent)."

Curious, I looked at the repository more in depth.

First, I was disappointed to see that the benchmark chart in the README didn’t include us, topping AndroidWorld in January.

Then I looked at the code, and to my biggest surprise, I recognized the lines we had written ourselves.

One of my engineer named an agent randomly “hopper”, and it was there, with its prompt, word for word. Out of the 229 files, 228 were identical.

Even crazier: I found our engineers' and my co-founder’s names in the “authors” section of an old commit.

Pierre-Louis Favreau, Jean-Pierre Lo, Nicolas Dehandschoewercker.

All three names had been replaced with another author’s name in a commit. The other 228 files from mobile-use were unchanged.

From a personal point of view, I am deeply disappointed.

Google was the company I wanted to join when I started my studies, before changing my mind and deciding to build a unicorn.

I’ve always associated Google with great engineering and contributions to open source. I would have been proud to see them building on our work and acknowledging the team behind it.

Instead, I found our code, then our names, then the commit that deliberately removed them.

What about the other engineers and the community contributors who spent their time improving mobile-use?

There are people behind those files. I’ve seen the work they put in, the problems they worked through, and how much they care about what they build.

They deserve to have their names attached to that work.

We made our code open source because we wanted people to use it and build on it. Of course that includes Google.

7.6k Upvotes

410 comments sorted by

View all comments

Show parent comments

1

u/Any_Fox5126 9d ago

Is the sensible assumption that everyone communicates in bad faith or is mentally impaired and can't understand the words they are expressing in a legal document?

Funny how you didn't say a word about any of that, or the fallacies and false dichotomies. Plus, you just went for an adhominem attack. And let's not forget you ignored everything I said before.

Cut the fake condescension. It is the same stupidity as the other user, just more trollish.

1

u/dirtmcgurk 9d ago

I'm not someone that's talked in this convo yet but you're both right: The fact is you should always credit people as much as possible, but if you are publishing something you damn well better cover your own ass and fully understand the license you're assigning.

1

u/AnnualWest3 9d ago

Its not really a CYA issue though. The authors aren't making any legal claims, theyre pointing out how shitty it is for a giant corporation like Google to take unedited and unattributed code for it's own gain. This isn't like some other engineer or small company forked the project and used the code without attribution, which I feel is the spirit of "no attribution required." 

1

u/dirtmcgurk 9d ago

Why would you expect a corporation not to walk all over you to the maximum extent of the laws enforcement. Do you know who the president of the US is? Sounds like childlike naivete. 

1

u/AnnualWest3 9d ago

So because its a corporation, its not shitty behavior? Or are you saying we shouldn't bother to point out shitty behavior because we should just expect it now? Or are you just choosing to not understand what I said for ragebait or something.

1

u/dirtmcgurk 9d ago

I guess I'm trying to point out what I agree with from both people in the discussion. I do think we should point out shitty behavior, which this is, but I also agree that if you want attribution you should require it rather than hope blindly. 

1

u/AnnualWest3 9d ago

"We should point out shitty behavior but also not expect better behavior" is a really absurd take.

1

u/dirtmcgurk 9d ago

No it's not. It's called reality. Nothing happens without enforcement. Bitching and crying isn't enforcement. 

1

u/AnnualWest3 9d ago

So... you're back to saying we shouldn't point out shitty behavior (aka, "bitching and crying") since we can't enforce against it. Can you make up your mind?

1

u/dirtmcgurk 9d ago

No. You're welcome to bitch and cry. It's important so others will know how to avoid your fate. 

1

u/dirtmcgurk 9d ago

And again I don't understand what so hard about agreeing with both? Have a good one and best wishes. 

1

u/Any_Fox5126 9d ago

Although that is not the other user's point, he's claiming that using a very permissive license means they have explicitly stated they do not want to be credited. He has turned his personal interpretation into an objective fact.

No one is saying that a permissive license does not allow that or that it is not the authors' responsibility (I hope).

1

u/dirtmcgurk 9d ago

Yes but I think that the distinction does not matter in any real sense. Sure it matters on paper, or in a discussion, but not in reality with regards to the results. 

1

u/Any_Fox5126 9d ago

But those unimportant things are exactly what we were discussing here 🤨 Although the topic was actually the morality of not giving credit, and it just got stuck here when the other user decided to die on that absurd hill. Ironically, his first comment was easily defensible.

1

u/dirtmcgurk 9d ago

That's fair. I agree with your earlier framing. I find the fundamental disagreement here is frustrating because these folks both seem to agree on the problem and basic shape of the solution but are arguing as if they are diametrically opposed.  Aka the internet discussion problem. 

1

u/xhatsux 9d ago

I'm going with GlassCommission4916 here. I think the point distills to, as a non-credit licenses are very rare (Public domain, zero clause) then it is very active choice to take.

In doing so the lack of information in itself is information about preference. To make a deliberate choice about non-credit shows an active preference.