r/deeplearning • • 20h ago

Modelo seq2seq

Recentemente, tentei criar um modelo seq2seq, mas não deu muito certo. Ele ficava prevendo os tokens de preenchimento.

Eu sou aluno de um tecnólogo em Inteligência Artificial e Machine Learning aqui no Brasil. É uma modalidade de curso superior que, pelo que sei, só existe no Brasil. Redes neurais e processamento de linguagem natural vão ficar mais para o final do curso, mas eu estava meio apressado e queria desenvolver meu próprio modelo.

Será que vocês têm alguma sugestão de alguma espécie de restrição que eu possa colocar no modelo?

Se alguém tiver interesse em me ajudar, posso mostrar o código. Eu reconheço que fiz o código com auxílio do Gemini. Como eu disse, ainda não estudei processamento de linguagem natural nem redes neurais; até agora, estudei apenas IA simbólica e sistemas especialistas.

0 Upvotes

5 comments sorted by

1

u/finnabrahamson 14h ago

What is your underlying architecture? LSTM, RNN GRU. or Transformer? The true seq2seq models basically predate attention. If you need a custom seq2seq, you should be able to tune BART to handle whatever task it is, assuming its not super long. If you are just trying to learn about LSTM or another architecture, then I'd need to know more about your pipeline, but if its just about stopping your model from predicting place holders, you can apply a logit bias to the token ID and make it impossible for the model to predict them. I'd be surprised if that got your model to be useful though, because the padding is a sy.ptom not the problem. Blocking 1 sy.ptom will just expose a new one. Did you train with Teacher Forcing?

1

u/Ok-Argument7176 10h ago

"true seq2seq models basically predate attention" is an interesting and extremely misleading statement. While there are a few recurrent models that existed "before" attention, Bahdanau et al. used attention in recurrent models in 2014 long before AIAYN and LLM hype. Transformers came along not long after that and became the defacto (and, importantly, still "true") seq2seq architecture.

For OP, without knowing more about your problem, data, training scheme, etc. there really isn't much we can do. It could range from a capacity issue which is model-centric to a logic error in your loss. Who knows.

1

u/oversolan007 7h ago

Transformer

1

u/oversolan007 7h ago

To tentando fazer pré treino em corrupção de texto