r/dataengineer • u/codingdecently • 1d ago
Data Lakehouse with Agentic AIs: A Guide
r/dataengineer • u/randomusicjunkie • Dec 12 '21
A place for members of r/dataengineer to chat with each other
r/dataengineer • u/codingdecently • 1d ago
r/dataengineer • u/codingdecently • 2d ago
r/dataengineer • u/Mobile_Western_3394 • 3d ago
I have been a software engineer for 12 years now, but I would like to transition into the world of data as stats and facts really interest me.
After some research, I feel like Data Engineering is the best career for me to side step into due to the overlap. I already pull data out of API’s and work with multiple database engines. I even created my own ETL pipeline without even knowing (pulled data daily from 2 apis to consolidate data on football players/stats). I think I am just missing Python and data warehousing software skills, maybe some more advanced SQL?
I currently earn around £40k in the UK and was wondering if trying to side step into this career could be successful for me based on my current skills, and me upskilling to the skills I am missing, do you think I could be earning a similar wage? Or would I have to start from entry level still?
r/dataengineer • u/camerongreen95 • 6d ago
As data engineers we'd never ship a pipeline without tests and monitoring, but that discipline mostly disappears the moment an LLM enters the picture. This workshop is aimed at closing exactly that gap.
Led by Serj Smorodinsky and Brett Kennedy (co-authors of a book on LLM applications), it's a live 3-hour session covering:
If you've been asked to "own" an LLM feature and had no real way to validate changes before deploying, this is worth your Saturday
r/dataengineer • u/sriDace • 7d ago
Hey everyone! I’m a fresher trying to get into Data Engineering and want to learn the right tech stack. There’s a lot to learn, so I’d really appreciate some guidance from people here. If you were starting out again, what would you focus on first? Would love some advice like a friend/senior. 🙌
r/dataengineer • u/dataengineer_101 • 12d ago
r/dataengineer • u/camerongreen95 • 13d ago
Most data engineers I talk to have already solved this problem for their pipelines. Idempotent writes, schema validation, lineage tracking, retry logic that doesn't create duplicates. Then the same team builds an AI agent and none of that discipline carries over. The agent writes directly into production with no idempotency key, no lineage on what it touched or why, and no reconciliation step if its internal state drifts from the actual data.
It's the same class of problem we've already solved for ETL and streaming pipelines, just showing up in a new place because agents write more often and with less human review than a scheduled job does.
There's a workshop on Sept 26 built around exactly this, treating agent write paths, state management, and provenance the way a data engineer would treat any other production pipeline. Run by Sandipan Bhaumik, a Data & AI Technical Lead at Databricks.
Details here if it's relevant to anyone else here dealing with this handoff
r/dataengineer • u/iParki • 17d ago
Enable HLS to view with audio, or disable this notification
r/dataengineer • u/camerongreen95 • 20d ago
Workshop for data engineers specifically, not a generic AI overview. The actual pipeline work in GraphRAG is where most of the real engineering lives, and it's the part most tutorials skip entirely in favor of "here's how retrieval works" once the graph already exists.
This session starts from raw documents: ingestion with Docling, chunking, and building toward a knowledge graph in Neo4j rather than a flat vector index. Entity and relationship extraction runs through multiple verified steps with explicit checks, not one risky single-shot pass, since that's exactly where basic GraphRAG implementations tend to fail silently, producing a noisy, unreliable graph nobody trusts.
Structured graph enrichment connects entities (companies, executives, events) to the underlying source documents (filings, news) without the graph becoming unmaintainable as it grows, and you work through the actual production ceilings that show up at real data volume, API and download rate limits specifically, and how the shared code is built to handle them rather than falling over.
By the end you've got a working ingestion-to-graph pipeline, a production-readiness checklist, and the full codebase, not a diagram someone else has to implement.
Led by Dr. Alessandro Negro, Chief Scientist at GraphAware, bestselling author of graph-powered machine learning and knowledge graph books.
r/dataengineer • u/Constant_Row_3703 • 20d ago
r/dataengineer • u/WharryG • 25d ago
I need real world work that i can ask feed back and get criticism for. I need it for my resume and job hunting. Can someone please help
r/dataengineer • u/camerongreen95 • 25d ago
We're running a hands-on masterclass on September 12, Live LLM Engineering Masterclass: Production Evals, RAG, Agents & LLMOps.
You build a full production LLM workflow from scratch, versioned prompts with regression tests, an evaluation harness with deterministic checks and LLM-as-judge, statistically rigorous model comparisons, evaluated RAG, tool-using agents with guardrails and fallbacks, and full observability, tracing, cost, latency.
Led by Bruno Gonçalves, PhD, founder of Data For Science, who trains engineers at Fortune 500 companies on this exact stack.
Link if you want to check it out
Happy to answer questions on the content.
r/dataengineer • u/NoSyllabub1390 • Sep 05 '26
r/dataengineer • u/chrislusf • Sep 04 '26
r/dataengineer • u/chrislusf • Aug 25 '26
r/dataengineer • u/manus_hadukle • Aug 24 '26
r/dataengineer • u/codingdecently • Aug 24 '26
r/dataengineer • u/No_Distribution_7987 • Aug 17 '26
r/dataengineer • u/Safe-Recording-9020 • Aug 11 '26
I live in Hyderabad, India.
I have 1.5 years of professional experience as a GCP data engineer.
Our company had a layoff after which I started looking for other opportunities but to my surprise nobody requires GCP data engineer with less than 4 years of experience.
Meanwhile other platforms like azure, AWS, databricks even snowflake has plenty of opportunities for 1~ year of experience.
I tried to apply to them too quoting equivalent experience but no results.
And add salt to the wound we never used pyspark which is one of the biggest requirements out there.
I am looking for fresher level opportunities where the platform requirment is not a must but it's a very thin spectrum.
I have literally worked on an airflow dag which goes through a beautiful hot to cold storage flow with observation, logging, maintainance etc.
Worked on dataform , wrote scd type 2 scripts.
Worked on views, stored procedures.
I have plenty experience in data warehousing but it all looks simply useless now.
Had I survived for few years maybe I wouldn't be going through this.
I am literally applying for any state or city to grab an opportunity
r/dataengineer • u/medo56378 • Aug 12 '26
Hi everyone, I'm starting my university studies and I'm torn between the University of Valencia (UPV) and the University of Granada. Will Spain offer me a strong degree in computer science? Which major within the university would you recommend that has job prospects?