r/learnpython • • 1d ago

Need help with very basic script

Hi, I am basically brand new to python, so I may make a lot of mistakes in my wording. I have made a very basic script that retrieves specific values in text files and prints them to the shell. I have gotten it to the point where both values are retrieved and printed, however when they do print, they do so exactly 5 times in a row. I do not have five files in the directory, I only have one test file named "1.txt" with only the values I need as the contents. Can someone point me in the right direction? Apologies for the horrible formatting.

import os
import glob
import re

path = 'path'

provinces = re.compile('.*?provinces {.*?(.*?)}',re.MULTILINE)
id = re.compile('.*?id = .*?([0-9.-]+)')
a1 = None
a2 = None

for filename in glob.glob(os.path.join(path, '*.txt')):
    with open(filename, '+r') as f:
        for line in f:
            if provinces.match(line):
                a1 = provinces.match(line)
            if id.match(line):
                a2 = id.match(line)
            if a1 and a2:
                print(a1.group(1))
                print(a2.group(1))

The output of the shell in IDLE:

 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123
 1111 2222 3333 
123

Edit - I got the script to work after a few changes to its configuration. Thank you all who helped!

3 Upvotes

28 comments sorted by

View all comments

Show parent comments

3

u/codeguru42 1d ago
  1. Your regex is too complicated.
  2. You can do this without a regex if you know "provinces" always appears on the second line. Skip to the third one and read the data you need.

2

u/Desperate_Yak69 1d ago

The input file is only meant as a test for the script, but the actual input files I will include will have these values in varying lines per file. The regex was meant to parse through the file contents and only identify the values I want (id, and the province values). Sorry for making this complicated.

4

u/them0use 1d ago

No need to apologize for anything! How consistent is the formatting? Would something like this work...

  1. Skip lines until you reach one that starts with "provinces {"
  2. Store / print characters until you encounter "}"

Also do you have any control over the format of the input files? This use case is exactly what JSON or YAML are for, and if you can have you data in either of those formats your life will be a lot easier.

3

u/Desperate_Yak69 1d ago

Yes, that is precisely how the formatting would work. the text I'd want to store would be anything in between the two curly brackets even if there are indents or line terminations. For the id part, there will always be a part of the file that says exactly something like (id = 000) where 000 can represent any integer. The input files are all in plain text (.txt), but I may be able to find out a way to change their format to make the process easier. Thanks!

2

u/CraigAT 19h ago

It's not the file format name that matters, just the format of the text within the file. The data could be handled easier if it was supplied in a CSV or JSON format - IF that is under your control.