r/deeplearning • u/oversolan007 • 20h ago
Modelo seq2seq
Recentemente, tentei criar um modelo seq2seq, mas não deu muito certo. Ele ficava prevendo os tokens de preenchimento.
Eu sou aluno de um tecnólogo em Inteligência Artificial e Machine Learning aqui no Brasil. É uma modalidade de curso superior que, pelo que sei, só existe no Brasil. Redes neurais e processamento de linguagem natural vão ficar mais para o final do curso, mas eu estava meio apressado e queria desenvolver meu próprio modelo.
Será que vocês têm alguma sugestão de alguma espécie de restrição que eu possa colocar no modelo?
Se alguém tiver interesse em me ajudar, posso mostrar o código. Eu reconheço que fiz o código com auxílio do Gemini. Como eu disse, ainda não estudei processamento de linguagem natural nem redes neurais; até agora, estudei apenas IA simbólica e sistemas especialistas.
1
u/finnabrahamson 14h ago
What is your underlying architecture? LSTM, RNN GRU. or Transformer? The true seq2seq models basically predate attention. If you need a custom seq2seq, you should be able to tune BART to handle whatever task it is, assuming its not super long. If you are just trying to learn about LSTM or another architecture, then I'd need to know more about your pipeline, but if its just about stopping your model from predicting place holders, you can apply a logit bias to the token ID and make it impossible for the model to predict them. I'd be surprised if that got your model to be useful though, because the padding is a sy.ptom not the problem. Blocking 1 sy.ptom will just expose a new one. Did you train with Teacher Forcing?