Ipython (DONE)
import pandas as pd
# import pandas_datareader.data as web
print(pd.__version__)0.25.0
Indexing
- which columns are index
- how to get column names
Type inference and data conversion
- user-defined value conversions
- custom missing values
Datetime parsing
- combine multiple columns into a single datetime info
Iterating
- support for iterating over chunks of very large files
Unclean data issues
- skip rows, footer, comments, thounsands separated by commas
reading function
- read_csv
- allow inference since csv has no data types
- read_fwf
- fixed width column format (no delimiters)
- read_clipboard
- read_pickle
- read_sql
- read result of SQL query (through SQLAlchemy) as pd.DataFrame
- read_csv
CSV¶
- read_csv
- header=None if there is no header
- names: specify column names
- index_col: to specify which col is the index
- one can specify multiple indices
- sep: can be a char (e.g., “,”) or a regular expression (e.g., “\s+”)
- skiprows: to skip certain rows known to store garbage
- na_values: allow to specify “sentinel” value for nan
- one can specify a dict to specify different nan for different columns
- comment
- parse_dates: try to parse all column with dates
- converters: function to be applied to columns to transform data
- nrows: max num of rows to read
- skip_footer: skip lines at the end
- encoding
# chunker = pd.read_csv("example/ex6.csv", chunksize=1000)
# returns an object that can be iterated onJSON¶
Binary format¶
- obj.to_pickle()
- pd.read_pickle()
HDF5¶
HDF = Hierarchical Data Format
store large quantities of scientific array data
Can store multiple datasets
Save metadata
Support compression
fixed format
- slower
- supports query operations
Web APIs¶
- Many websites have public API providing data feed via JSON or other format.
requestslibrary
import requests
# Get last 30 GitHub issues from pandas.
url = "https://api.github.com/repos/pandas-dev/pandas/issues"
resp = requests.get(url)resp<Response [200]>data = resp.json()
data[0]["title"]'DEPR: Deprecate sort=None for union and implement sort=True'