Chatistics – Turn Telegram, WhatsApp, Messenger Chatlogs to DataFrames

栏目: IT技术 · 发布时间: 4年前

内容简介:Python 3 scripts to convert chat logs from various messaging platforms into Panda DataFrames.Can also generate histograms and word clouds from the chat logs.10 Jan 2020:UPDATED

Chatistics

Python 3 scripts to convert chat logs from various messaging platforms into Panda DataFrames.Can also generate histograms and word clouds from the chat logs.

Chatistics – Turn Telegram, WhatsApp, Messenger Chatlogs to DataFrames Chatistics – Turn Telegram, WhatsApp, Messenger Chatlogs to DataFrames

Changelog

10 Jan 2020:UPDATED ALL THE THINGS! Thanks to mar-muel and manueth , pretty much everything has been updated and improved, and WhatsApp is now supported!

21 Oct 2018:Updated Facebook Messenger and Google Hangouts parsers to make them work with the new exported file formats.

9 Feb 2018:Telegram support added thanks to bmwant .

24 Oct 2016:Initial release supporting Facebook Messenger and Google Hangouts.

Support Matrix

Platform Direct Chat Group Chat
Facebook Messenger
Google Hangouts
Telegram
WhatsApp

Exported data

Data exported for each message regardless of the platform:

Column Content
timestamp UNIX timestamp (in seconds)
conversationId A conversation ID, unique by platform
conversationWithName Name of the other people in a direct conversation, or name of the group conversation
senderName Name of the sender
outgoing Boolean value whether the message is outgoing/coming from owner
text Text of the message
language Language of the conversation as inferred by langdetect
platform Platform (see support matrix above)

Exporting your chat logs

1. Download your chat logs

Google Hangouts

Warning:Google Hangouts archives can take a long time to be ready for download - up to one hour in our experience.

Hangouts.json
./raw_data/hangouts/

Facebook Messenger

Warning:Facebook archives can take a very long time to be ready for download - up to 12 hours! They can weight several gigabytes. Start with an archive containing just a few months of data if you want to quickly get started, this shouldn’t take more than a few minutes to complete.

  1. Go to the page “Your Facebook Information”: https://www.facebook.com/settings?tab=your_facebook_information
  2. Click on “Download Your Information”
  3. Select the date range you want. The format must be JSON. Media won’t be used, so you can set the quality to “Low” to speed things up.
  4. Click on “Deselect All”, then scroll down to select “Messages” only
  5. Click on “Create File” at the top of the list. It will take Facebook a while to generate your archive.
  6. Once the archive is ready, download and extract it, then move the content of the messages folder into ./raw_data/messenger/

WhatsApp

Unfortunately, WhatsApp only lets you export your conversations from your phone and one by one .

  1. On your phone, open the chat conversation you want to export
  2. On Android , tap on > More > Export chat . On iOS , tap on the interlocutor’s name > Export chat
  3. Choose “Without Media”
  4. Send chat to yourself eg via Email
  5. Unpack the archive and add the individual .txt files to the folder ./raw_data/whatsapp/

Telegram

The Telegram API works differently: you will first need to setup Chatistics, then query your chat logs programmatically. This process is documented below. Exporting Telegram chat logs is very fast.

2. Setup Chatistics

First, install the required Python packages:

Use conda (recommended)

conda env create -f environment.yml
conda activate chatistics

Or virtualenv

virtualenv chatistics
source chatistics/bin/activate
pip install -r requirements.txt

You can now parse the messages by using the command python parse.py <platform> <arguments> .

By default the parsers will try to infer your own name (i.e. your username) from the data. If this fails you can provide your own name to the parser by providing the --own-name argument. The name should match your name exactly as used on that chat platform.

# Google Hangouts
python parse.py hangouts

# Facebook Messenger
python parse.py messenger

# WhatsApp
python parse.py whatsapp

Telegram

  1. Create your Telegram application to access chat logs ( instructions ). You will need api_id and api_hash which we will now set as environment variables.
  2. Run cp secrets.sh.example secrets.sh and fill in the values for the environment variables TELEGRAM_API_ID , TELEGRAMP_API_HASH and TELEGRAM_PHONE (your phone number including country code).
  3. Run source secrets.sh
  4. Execute the parser script using python parse.py telegram

The pickle files will now be ready for analysis in the data folder!

For more options use the -h argument on the parsers (e.g. python parse.py telegram --help ).

3. All done! Play with your data

Chatistics can print the chat logs as raw text. It can also create histograms, showing how many messages each interlocutor sent, or generate word clouds based on word density and a base image.

Export

You can view the data in stdout (default) or export it to csv, json, or as a Dataframe pickle.

python export.py

You can use the same filter options as described above in combination with an output format option:

-f {stdout,json,csv,pkl}, --format {stdout,json,csv,pkl}
                        Output format (default: stdout)

Histograms

Plot all messages with:

python visualize.py breakdown

Among other options you can filter messages as needed (also see python visualize.py breakdown --help ):

--platforms {telegram,whatsapp,messenger,hangouts}
                        Use data only from certain platforms (default: ['telegram', 'whatsapp', 'messenger', 'hangouts'])
  --filter-conversation
                        Limit by conversations with this person/group (default: [])
  --filter-sender
                        Limit to messages sent by this person/group (default: [])
  --remove-conversation
                        Remove messages by these senders/groups (default: [])
  --remove-sender
                        Remove all messages by this sender (default: [])
  --contains-keyword
                        Filter by messages which contain certain keywords (default: [])
  --outgoing-only       
                        Limit by outgoing messages (default: False)
  --incoming-only       
                        Limit by incoming messages (default: False)

Eg to see all the messages sent between you and Jane Doe:

python visualize.py breakdown --filter-conversation "Jane Doe"

To see the messages sent to you by the top 10 people with whom you talk the most:

python visualize.py breakdown -n 10 --incoming-only

Chatistics – Turn Telegram, WhatsApp, Messenger Chatlogs to DataFrames

You can also plot the conversation densities using the --as-density flag.

Chatistics – Turn Telegram, WhatsApp, Messenger Chatlogs to DataFrames

Word Cloud

You will need a mask file to render the word cloud. The white bits of the image will be left empty, the rest will be filled with words using the color of the image. See the WordCloud library documentation for more information.

python visualize.py cloud -m raw_outlines/users.jpg

You can filter which messages to use using the same flags as with histograms.

Development

Install dev environment using

conda env create -f environment_dev.yml

Run tests from project root using

python -m pytest

Improvement ideas

  • Parsers for more chat platforms: Signal? Pidgin? …
  • Handle group chats on more platforms.
  • See open issues for more ideas.

Pull requests are welcome!

Misc

  • Word cloud generated using https://github.com/amueller/word_cloud
  • Stopwords from https://github.com/6/stopwords-json
  • Code under MIT license

以上就是本文的全部内容,希望对大家的学习有所帮助,也希望大家多多支持 码农网

查看所有标签

猜你喜欢:

本站部分资源来源于网络,本站转载出于传递更多信息之目的,版权归原作者或者来源机构所有,如转载稿涉及版权问题,请联系我们

阿里传

阿里传

波特·埃里斯曼 / 张光磊、吕靖纬、崔玉开 / 中信出版社 / 2015-9-15 / CNY 49.00

你只知道阿里巴巴故事的中国部分,而这本书会完整呈现故事的全部。 波特•埃里斯曼是阿里巴巴创业时期为数不多的外国高管。他于2000~2008年在阿里巴巴担任副总裁,这本书记录了他在阿里巴巴8年的时间里的创业故事、商业经验以及在阿里巴巴和马云、蔡崇信、关明生等阿里巴巴早期团队并肩奋战的故事。 在波特眼中,阿里巴巴的成功经验和模式是可以复制的,阿里巴巴曾经犯过的错误,走过的弯路,我们也可以绕......一起来看看 《阿里传》 这本书的介绍吧!

RGB转16进制工具
RGB转16进制工具

RGB HEX 互转工具

图片转BASE64编码
图片转BASE64编码

在线图片转Base64编码工具

HEX HSV 转换工具
HEX HSV 转换工具

HEX HSV 互换工具