

기계 번역으로 제공되는 번역입니다. 제공된 번역과 원본 영어의 내용이 상충하는 경우에는 영어 버전이 우선합니다.

# 테이블을 CSV 파일로 내보내기
<a name="examples-export-table-csv"></a>

이 Python 예제에서는 문서의 이미지에서 쉼표로 구분된 값(CSV) 파일로 테이블을 내보내는 방법을 보여줍니다.

동기식 문서 분석 예제는 [AnalyzeDocument](https://docs.aws.amazon.com/textract/latest/APIReference/API_AnalyzeDocument.html)에 대한 호출에서 테이블 정보를 수집합니다. 비동기 문서 분석 예제는 [StartDocumentAnalysis](https://docs.aws.amazon.com/textract/latest/APIReference/API_StartDocumentAnalysis.html)를 호출한 다음 [GetDocumentAnalysis](https://docs.aws.amazon.com/textract/latest/APIReference/API_GetDocumentAnalysis.html)의 결과를 `Block` 객체로 검색합니다.

테이블 정보는 [AnalyzeDocument](https://docs.aws.amazon.com/textract/latest/APIReference/API_AnalyzeDocument.html)에 대한 호출에서 [객체 차단](https://docs.aws.amazon.com/textract/latest/APIReference/API_Block.html)으로 반환됩니다. 자세한 내용은 [테이블](how-it-works-tables.md) 단원을 참조하십시오. `Block` 객체는 테이블 데이터를 CSV 파일로 내보내는 데 사용되는 맵 구조에 저장됩니다.

------
#### [ Synchronous ]

이 예제에서는 함수를 사용합니다.
+ `get_table_csv_results` - [AnalyzeDocument](https://docs.aws.amazon.com/textract/latest/APIReference/API_AnalyzeDocument.html)를 호출하고 문서에서 감지된 테이블 맵을 빌드합니다. 감지된 모든 테이블의 CSV 표현을 생성합니다.
+ `generate_table_csv` - 개별 테이블에 대한 CSV 파일을 생성합니다.
+ `get_rows_columns_map` - 맵에서 행과 열을 가져옵니다.
+ `get_text` - 셀에서 텍스트를 가져옵니다.

**테이블을 CSV 파일로 내보내려면**

1. 환경을 구성합니다. 자세한 내용은 [사전 조건](examples-blocks.md#examples-prerequisites) 단원을 참조하십시오.

1. 다음 예제 코드를 *textract\_python\_table\_parser.py*라는 파일에 저장합니다. 함수에서를 역할을 수임할 수 있는 프로파일의 `profile-name` 이름으로 바꾸고를 코드를 실행하려는 리전`region`으로 `get_table_csv_results`바꿉니다.

   ```
   import webbrowser, os
   import json
   import boto3
   import io
   from io import BytesIO
   import sys
   from pprint import pprint
   
   
   def get_rows_columns_map(table_result, blocks_map):
       rows = {}
       scores = []
       for relationship in table_result['Relationships']:
           if relationship['Type'] == 'CHILD':
               for child_id in relationship['Ids']:
                   cell = blocks_map[child_id]
                   if cell['BlockType'] == 'CELL':
                       row_index = cell['RowIndex']
                       col_index = cell['ColumnIndex']
                       if row_index not in rows:
                           # create new row
                           rows[row_index] = {}
                       
                       # get confidence score
                       scores.append(str(cell['Confidence']))
                           
                       # get the text value
                       rows[row_index][col_index] = get_text(cell, blocks_map)
       return rows, scores
   
   
   def get_text(result, blocks_map):
       text = ''
       if 'Relationships' in result:
           for relationship in result['Relationships']:
               if relationship['Type'] == 'CHILD':
                   for child_id in relationship['Ids']:
                       word = blocks_map[child_id]
                       if word['BlockType'] == 'WORD':
                           if "," in word['Text'] and word['Text'].replace(",", "").isnumeric():
                               text += '"' + word['Text'] + '"' + ' '
                           else:
                               text += word['Text'] + ' '
                       if word['BlockType'] == 'SELECTION_ELEMENT':
                           if word['SelectionStatus'] =='SELECTED':
                               text +=  'X '
       return text
   
   
   def get_table_csv_results(file_name):
   
       with open(file_name, 'rb') as file:
           img_test = file.read()
           bytes_test = bytearray(img_test)
           print('Image loaded', file_name)
   
       # process using image bytes
       # get the results
       session = boto3.Session(profile_name='profile-name')
       client = session.client('textract', region_name='region')
       response = client.analyze_document(Document={'Bytes': bytes_test}, FeatureTypes=['TABLES'])
   
       # Get the text blocks
       blocks=response['Blocks']
       pprint(blocks)
   
       blocks_map = {}
       table_blocks = []
       for block in blocks:
           blocks_map[block['Id']] = block
           if block['BlockType'] == "TABLE":
               table_blocks.append(block)
   
       if len(table_blocks) <= 0:
           return "<b> NO Table FOUND </b>"
   
       csv = ''
       for index, table in enumerate(table_blocks):
           csv += generate_table_csv(table, blocks_map, index +1)
           csv += '\n\n'
   
       return csv
   
   def generate_table_csv(table_result, blocks_map, table_index):
       rows, scores = get_rows_columns_map(table_result, blocks_map)
   
       table_id = 'Table_' + str(table_index)
       
       # get cells.
       csv = 'Table: {0}\n\n'.format(table_id)
   
       for row_index, cols in rows.items():
           for col_index, text in cols.items():
               col_indices = len(cols.items())
               csv += '{}'.format(text) + ","
           csv += '\n'
           
       csv += '\n\n Confidence Scores % (Table Cell) \n'
       cols_count = 0
       for score in scores:
           cols_count += 1
           csv += score + ","
           if cols_count == col_indices:
               csv += '\n'
               cols_count = 0
   
       csv += '\n\n\n'
       return csv
   
   def main(file_name):
       table_csv = get_table_csv_results(file_name)
   
       output_file = 'output.csv'
   
       # replace content
       with open(output_file, "wt") as fout:
           fout.write(table_csv)
   
       # show the results
       print('CSV OUTPUT FILE: ', output_file)
   
   
   if __name__ == "__main__":
       file_name = sys.argv[1]
       main(file_name)
   ```

1. 명령 프롬프트에 다음 명령을 입력합니다. 를 분석하려는 문서 이미지 파일의 이름으로 `file` 바꿉니다.

   ```
   python textract_python_table_parser.py {{file}}
   ```

예제를 실행하면 CSV 출력이 라는 파일에 저장됩니다`output.csv`.

------
#### [ Asynchronous ]

이 예제에서는 두 가지 스크립트를 사용합니다. 첫 번째 스크립트는를 사용하여 문서를 비동기식으로 분석`StartDocumentAnalysis`하는 프로세스를 시작하고에서 반환한 `Block` 정보를 가져옵니다`GetDocumentAnalysis`. 두 번째 스크립트는 각 페이지에 대해 반환된 `Block` 정보를 가져와서 데이터의 형식을 테이블로 지정하고 테이블을 CSV 파일에 저장합니다.

**테이블을 CSV 파일로 내보내려면**

1. 환경을 구성합니다. 자세한 내용은 [사전 조건](examples-blocks.md#examples-prerequisites) 단원을 참조하십시오.

1. 에 제공된 지침을 따랐는지 확인합니다. 단원을 참조하십시오[비동기 작업을 위한 Amazon Textract 구성](api-async-roles.md). 해당 페이지에 설명된 프로세스를 통해 비동기 작업의 완료 상태에 대한 메시지를 보내고 받을 수 있습니다.

1. 다음 코드 예제에서의 값을 2단계에서 생성한 역할에 할당된 `roleArn` Arn으로 바꿉니다. 의 값을 문서가 포함된 S3 버킷의 `bucket` 이름으로 바꿉니다. 의 값을 S3 버킷의 문서 `document` 이름으로 바꿉니다. 의 값을 버킷의 리전 `region_name` 이름으로 바꿉니다.

   다음 예제 코드를 *start\_doc\_analysis\_for\_table\_extraction.py라는 파일에 저장합니다*.

   ```
   import boto3
   import time
   
   class DocumentProcessor:
   
       jobId = ''
       region_name = ''
   
       roleArn = ''
       bucket = ''
       document = ''
   
       sqsQueueUrl = ''
       snsTopicArn = ''
       processType = ''
   
       def __init__(self, role, bucket, document, region):
           self.roleArn = role
           self.bucket = bucket
           self.document = document
           self.region_name = region
   
           self.textract = boto3.client('textract', region_name=self.region_name)
           self.sqs = boto3.client('sqs')
           self.sns = boto3.client('sns')
   
       def ProcessDocument(self):
   
           jobFound = False
   
           response = self.textract.start_document_analysis(DocumentLocation={'S3Object': {'Bucket': self.bucket, 'Name': self.document}},
                   FeatureTypes=["TABLES", "FORMS"], NotificationChannel={'RoleArn': self.roleArn, 'SNSTopicArn': self.snsTopicArn})
           print('Processing type: Analysis')
   
           print('Start Job Id: ' + response['JobId'])
   
           print('Done!')
   
       def CreateTopicandQueue(self):
   
           millis = str(int(round(time.time() * 1000)))
   
           # Create SNS topic
           snsTopicName = "AmazonTextractTopic" + millis
   
           topicResponse = self.sns.create_topic(Name=snsTopicName)
           self.snsTopicArn = topicResponse['TopicArn']
   
           # create SQS queue
           sqsQueueName = "AmazonTextractQueue" + millis
           self.sqs.create_queue(QueueName=sqsQueueName)
           self.sqsQueueUrl = self.sqs.get_queue_url(QueueName=sqsQueueName)['QueueUrl']
   
           attribs = self.sqs.get_queue_attributes(QueueUrl=self.sqsQueueUrl,
                                                   AttributeNames=['QueueArn'])['Attributes']
   
           sqsQueueArn = attribs['QueueArn']
   
           # Subscribe SQS queue to SNS topic
           self.sns.subscribe(TopicArn=self.snsTopicArn, Protocol='sqs', Endpoint=sqsQueueArn)
   
           # Authorize SNS to write SQS queue
           policy = """{{
         "Version":"2012-10-17",		 	 	 
         "Statement":[
           {{
             "Sid":"MyPolicy",
             "Effect":"Allow",
             "Principal" : {{"AWS" : "*"}},
             "Action":"SQS:SendMessage",
             "Resource": "{}",
             "Condition":{{
               "ArnEquals":{{
                 "aws:SourceArn": "{}"
               }}
             }}
           }}
         ]
       }}""".format(sqsQueueArn, self.snsTopicArn)
   
           response = self.sqs.set_queue_attributes(
               QueueUrl=self.sqsQueueUrl,
               Attributes={
                   'Policy': policy
               })
   
   def main():
       roleArn = 'role-arn'
       bucket = 'bucket-name'
       document = 'document-name'
       region_name = 'region-name'
   
       analyzer = DocumentProcessor(roleArn, bucket, document, region_name)
       analyzer.CreateTopicandQueue()
       analyzer.ProcessDocument()
   
   if __name__ == "__main__":
       main()
   ```

1. 코드를 실행합니다. 코드는 JobId를 인쇄합니다. 이 JobId를 복사합니다.

1.  작업이 처리를 완료할 때까지 기다린 후 다음 코드를 *get\_doc\_analysis\_for\_table\_extraction.py*라는 파일에 복사합니다. 의 값을 이전에 복사한 작업 ID`jobId`로 바꿉니다. 의 값을 Textract 역할과 연결된 리전의 `region_name` 이름으로 바꿉니다. 의 값을 출력 CSV`file_name`를 제공하려는 이름으로 바꿉니다.

   ```
   import boto3
   from pprint import pprint
   
   jobId = ''
   region_name = ''
   file_name = ''
   
   textract = boto3.client('textract', region_name=region_name)
   
   # Display information about a block
   def DisplayBlockInfo(block):
       print("Block Id: " + block['Id'])
       print("Type: " + block['BlockType'])
       if 'EntityTypes' in block:
           print('EntityTypes: {}'.format(block['EntityTypes']))
   
       if 'Text' in block:
           print("Text: " + block['Text'])
   
       if block['BlockType'] != 'PAGE':
           print("Confidence: " + "{:.2f}".format(block['Confidence']) + "%")
   
   def GetResults(jobId, file_name):
       maxResults = 1000
       paginationToken = None
       finished = False
   
       while finished == False:
   
           response = None
   
           if paginationToken == None:
               response = textract.get_document_analysis(JobId=jobId, MaxResults=maxResults)
           else:
               response = textract.get_document_analysis(JobId=jobId, MaxResults=maxResults,
                                                              NextToken=paginationToken)
   
           blocks = response['Blocks']
           table_csv = get_table_csv_results(blocks)
           output_file = file_name + ".csv"
           # replace content
           with open(output_file, "at") as fout:
               fout.write(table_csv)
           # show the results
           print('Detected Document Text')
           print('Pages: {}'.format(response['DocumentMetadata']['Pages']))
           print('OUTPUT TO CSV FILE: ', output_file)
   
           # Display block information
           for block in blocks:
               DisplayBlockInfo(block)
               print()
               print()
   
           if 'NextToken' in response:
               paginationToken = response['NextToken']
           else:
               finished = True
   
   
   def get_rows_columns_map(table_result, blocks_map):
       rows = {}
       for relationship in table_result['Relationships']:
           if relationship['Type'] == 'CHILD':
               for child_id in relationship['Ids']:
                   try:
                       cell = blocks_map[child_id]
                       if cell['BlockType'] == 'CELL':
                           row_index = cell['RowIndex']
                           col_index = cell['ColumnIndex']
                           if row_index not in rows:
                               # create new row
                               rows[row_index] = {}
   
                           # get the text value
                           rows[row_index][col_index] = get_text(cell, blocks_map)
                   except KeyError:
                       print("Error extracting Table data - {}:".format(KeyError))
                       pass
       return rows
   
   
   def get_text(result, blocks_map):
       text = ''
       if 'Relationships' in result:
           for relationship in result['Relationships']:
               if relationship['Type'] == 'CHILD':
                   for child_id in relationship['Ids']:
                       try:
                           word = blocks_map[child_id]
                           if word['BlockType'] == 'WORD':
                               text += word['Text'] + ' '
                           if word['BlockType'] == 'SELECTION_ELEMENT':
                               if word['SelectionStatus'] == 'SELECTED':
                                   text += 'X '
                       except KeyError:
                           print("Error extracting Table data - {}:".format(KeyError))
   
       return text
   
   
   def get_table_csv_results(blocks):
   
       pprint(blocks)
   
       blocks_map = {}
       table_blocks = []
       for block in blocks:
           blocks_map[block['Id']] = block
           if block['BlockType'] == "TABLE":
               table_blocks.append(block)
   
       if len(table_blocks) <= 0:
           return "<b> NO Table FOUND </b>"
   
       csv = ''
       for index, table in enumerate(table_blocks):
           csv += generate_table_csv(table, blocks_map, index + 1)
           csv += '\n\n'
           # In order to generate separate CSV file for every table, uncomment code below
           #inner_csv = ''
           #inner_csv += generate_table_csv(table, blocks_map, index + 1)
           #inner_csv += '\n\n'
           #output_file = file_name + "___" + str(index) + ".csv"
           # replace content
           #with open(output_file, "at") as fout:
           #    fout.write(inner_csv)
   
       return csv
   
   
   def generate_table_csv(table_result, blocks_map, table_index):
       rows = get_rows_columns_map(table_result, blocks_map)
   
       table_id = 'Table_' + str(table_index)
   
       # get cells.
       csv = 'Table: {0}\n\n'.format(table_id)
   
       for row_index, cols in rows.items():
   
           for col_index, text in cols.items():
               csv += '{}'.format(text) + ","
           csv += '\n'
   
       csv += '\n\n\n'
       return csv
   
   response_blocks = GetResults(jobId, file_name)
   ```

1. 코드를 실행합니다.

   결과를 얻은 후에는 연결된 SNS 및 SQS 리소스를 삭제해야 합니다. 그렇지 않으면 해당 리소스에 대한 요금이 발생할 수 있습니다.

------