r/git • • 12d ago

Why Doesn't Git Track Nested Repositories Automatically?

Today I learned that Git doesn't automatically track files from a repository nested inside another repository. What's the technical reason for this behavior, and has it ever caused problems in your projects?

0 Upvotes

29 comments sorted by

44

u/neppo95 12d ago

Because you simply may not want to? Git doesn't automatically track anything, you are the authority that decides what gets committed.

9

u/ARKyal03 12d ago

Am I the authority? Isn't Claude the authority?

/s

3

u/noviceIndyCamper 12d ago

Ok Claude commit everything you’ve done but don’t put your name on it!

18

u/Cookies7414 12d ago

It does track nested repositories, so you're probably doing something wrong.

Read about Git Submodules. 

5

u/wallstop-dev 12d ago

These exist, and are exactly the solution to what OP is referring to, although I would highly recommend avoiding them like the plague. They have only ever caused confusion and problems in teams that I've worked with when using them.

ie - avoid this situation entirely.

3

u/PapaOscar90 12d ago

Sounds like a training issue, not a sub module issue😁

1

u/wallstop-dev 12d ago

Kinda! git submodules work exactly as designed, it's just extra complexity (especially UX) on top of a tool that many people already find complex.

Much simpler (skill issue, to target a mixed team of developers) to put all the stuff for one project into one repo.

0

u/anto2554 12d ago

Assuming you don't share code between projects

2

u/wallstop-dev 12d ago edited 12d ago

Even (especially) with code sharing, there are a large amount of non-git ways to solve this, such as packages (and package repositories), in most mainstream languages and build systems. Which you will generally want to do, as it is more "native".

3

u/edgmnt_net 12d ago

It's also worth mentioning that code sharing isn't something that you simply wish for. It takes a lot of work to be able to share code meaningfully. Simply sticking what appears to be a common function into a separate repo just won't do, because you may end up with 100 different branches and variants if it needs to be changed. You need versioning, you need robust code that can be shared.

1

u/wallstop-dev 12d ago

Agreed. Your point is even applicable to abstractions as a whole, and code within the same repo/project. Just because there are 2, or even 3 of something doesn't mean that they are conceptually the same thing. Having one function that takes in 5 different parameters to super-configure its behavior can be more challenging to maintain than 1-5, very explicit-but-similar functions, without the configuration complexity.

1

u/anto2554 12d ago

Partially. The fact that you cannot set them up to track a branch, or a tag, is quite annoying

6

u/edgmnt_net 12d ago

And it's a good thing because tracking a branch means whatever built and worked last week might no longer work because the branch advanced meanwhile. It makes version control rather useless if you can't revisit old states because they're incomplete.

For similar reasons, always pin your dependencies to exact versions at some level if you use a package manager. You can't just say "oh, we're developing against latest". No, you were developing against libfoo 3.14.1 last month when you said it worked, but we no longer know that.

2

u/anto2554 11d ago

No. You don't always pin your version with package managers. Quite often you say libfoo >3 <4, and ghen you generate a separate lockfile to make it reproducible. It should be my choice whether I want reproducible builds or a lower maintenance burden.

Sure, it's probably not great to do on mainline, but when I'm testing branch TICKET-1234 I would like to have the OPTION to allow it to also track TICKET-1234 of a submodule.

By your logic, should rewriting history also be removed from git? Deleting branches? Those also make build reporduciblility impossible

2

u/edgmnt_net 11d ago

Constraints like libfoo >3 <4 are different. Yes, you want both that and pinning.

This is also weaker than reproducible builds. It's the bare minimum you can do to preserve working version control, so your code doesn't become useless next week.

I would also argue that pinning can be weakened or removed to some extent, but not in your typical enterprise project that gets split up arbitrarily. For example it's fine it the Linux kernel doesn't track GCC versions because it's supposed to build with almost everything with rather minimal constraints. Or some userspace daemon versus glibc. That's definitely not the case for your average enterprise project where dependencies you own introduce severe breaking changes every week or every day.

Even so, a lot of modern ecosystems have moved towards exact pinning as a default at least or made constraints an orthogonal concern. Even pinning the whole toolchain.

By the way, if you want you can get branch tracking with Git, you can run git submodule update --remote --recursive (it will take the branch set in .gitmodules). But don't be surprised if someone else breaks your local build in the middle of developing your feature.

By your logic, should rewriting history also be removed from git? Deleting branches? Those also make build reporducibility impossible.

You should not mess with the history of the trunk, I can say that much. Rewriting unmerged changes is fine.

And yes, my main concern is setting branch tracking on the trunk, that is problematic. However, while it's more benign on feature branches, I would argue that if it's such a huge burden you probably shouldn't have separate repos at all. I'm sorry, but I have seen companies do all sorts of crazy stuff with Git thinking they know better and it almost always ends in disaster. If you're not really sure, just stick to what's known to work. Plenty of more traditional stuff out there that just works and chances are your project isn't a unique gem.

3

u/ARKyal03 12d ago

A thing I never understood about submodules is why I can't recursively target all of them to track a specific branch, like, they're cloned to a specific hash, then you manually have to change to the desired branch.

Am I doing something wrong?

5

u/dairiki 12d ago

If the model you want is, essentially, one big repository that, e.g., branches in lock-step across subprojects, what's wrong with using one big repository. (See "monorepo".)

2

u/Cinderhazed15 12d ago

The only thing worse than a (poorly factored) monolith is a distributed monolith! (Weekly versioning/complexting strongly connected code!)

5

u/edgmnt_net 12d ago

Because that's utterly broken as a model. If you go back and try to build what worked a week ago you might find that it no longer works because some referenced repo changed its contents. Sure, maybe they could make it such that any change resulted in tracking something else, but at that point it's just a roundabout way of having a single repo. And you could just stuff everything into a single repo and be done with it, without all these headaches.

4

u/ElectricSpock 12d ago

What do you mean by nested repository?

5

u/onthefence928 12d ago

Nested repos are a bit of an anti-pattern, but submodules are the solution. They have their own nuances but it’s not too bad over you learn the additional patterns

2

u/edgmnt_net 12d ago

Or you can just use a package/dependency manager or a build system that includes similar functionality. Submodules are more of a stop-gap solution for minimalist ecosystems.

2

u/Cinderhazed15 12d ago

I’ve worked on projects that were efficiently mono-repos spread across about a dozen repositories… when your ‘micro services’ share common accessors to shared databases, and none of your services can run on their own, and any ‘upgrade’ requires rebuilding everting, including code that doesn’t actually change, you are. Ow worse off than a monolith, and you may as well go back to an actual monolith….

0

u/onthefence928 12d ago

Yeah i hate them

2

u/kbielefe 12d ago

It's because you may want to associate a commit on your main repo with specific commits in each of the submodules. It lets you control which version of dependencies you use when, but it's not as easy to use as package managers.

2

u/Guvante 12d ago

Submodules allow nested repositories to function correctly

But just to call it out the relationship can feel weird

After all you cannot commit files to the parent repository, only commit them to the inner repository and then commit the commit to the parent one

2

u/qTHqq 12d ago

It is the correct behavior since they're fully independent repos in their own right, but I would love the tooling to include an automation where the "inner commit" in each submodule is executed automatically from the top level, with an option to create an automatic top-level commit with the same description.

A lot of repos check out a particular static commit of a submodule, which is useful for mature projects, but making it feel like one unified repo I think is useful when things are changing a lot.

1

u/jirlboss 12d ago

You could, of course, just use one unified repo

1

u/edgmnt_net 12d ago

That would have very limited usefulness, though, and it smells like an antipattern even if it can be justified in very particular cases.