The migbirdHIP Workflow
migbirdHIP-workflow.RmdTable of Contents
Introduction
The migbirdHIP package was created by the U.S. Fish and
Wildlife Service (USFWS) to process, clean, and visualize Harvest
Information Program (HIP) registration data.
Installation
The package can be installed from the USFWS GitHub repository using:
pak::pak("USFWS/migbirdHIP")Load
Load migbirdHIP after installation. The package will
return a startup message with the version you have installed and what
season of HIP data the package version was intended for.
## migbirdHIP v2026.0.2
## Compatible with 2026-2027 HIP data
Functions overview
The flowchart below is a visual guide to the order in which functions
are used. Some functions are only used situationally and some issues
with the data cannot be solved using a function at all. The general
process of handling HIP data is demonstrated in the flowchart; every
exported function in the migbirdHIP package is included in
it and the vignette.
Overview of migbirdHIP functions in a flowchart format.
Example data
Example data are included in the migbirdHIP R package in
two formats. Raw data files are provided with fake names, addresses, and
other registration values, mainly to include examples and tests for
reading in data. In addition, small data objects derived from the fake
registrations are exported to users.
Raw data
Raw data in the form of fake HIP registration data files were
generated with data-raw/create_fake_HIP_data.R. The 49
files (one for each state) are stored in the migbirdHIP R
package under inst/extdata/. You can test read_hip() with the files in this directory.
Exported data
There are 7 data objects stored in /data/ as
.rda files. Users can call these objects by name (e.g.,
DF_TEST_MINI) and query their documentation (e.g.,
?DF_TEST_MINI) from the console after loading
migbirdHIP.
The biggest object is DF_TEST_MINI with ~1,600 rows.
This tibble represents “raw data” and doesn’t contain the
record_key field. The other 6 objects all begin with the
DF_TEST_TINI_ prefix and represent HIP data at various
points during pre-processing. They are small, containing only 3 rows per
object, but exhibit properties that are expected from each intermediate
output. For example, DF_TEST_TINI_PROOFED contains the
errors field. Users can run these objects in
migbirdHIP package functions.
List of exported data objects:
DF_TEST_MINIDF_TEST_TINI_READDF_TEST_TINI_CLEANEDDF_TEST_TINI_CURRENTDF_TEST_TINI_DEDUPEDDF_TEST_TINI_PROOFEDDF_TEST_TINI_CORRECTED
Part A: Data Import and Cleaning
fileRename
The fileRename() function renames non-standard file
names to the standard format. We expect a file name to be a 2-letter
capitalized state abbreviation followed by YYYYMMDD (indicating the date
data were submitted).
Some states submit files using a 5-digit file name format, containing
2 letters for the state abbreviation followed by a 3-digit Julian date.
To convert these 5-digit filenames to the standard format (a requirement
to read data properly with read_hip()), supply
fileRename() with the directory containing HIP data. File
names will be automatically overwritten with the YYYYMMDD format date
corresponding to the submitted Julian date.
This function also converts lowercase state abbreviations to uppercase.
The current hunting season year must be supplied to the
year parameter to accurately convert dates.
fileRename(path = "C:/HIP/raw_data/DL0901", year = 2026)fileCheck
Check if any files in the input folder have already been written to
the processed folder using fileCheck().
fileCheck(
raw_path = "C:/HIP/raw_data/DL0901/",
processed_path = "C:/HIP/corrected_data/"
)read_hip
Read HIP data from fixed-width .txt files using
read_hip(). Files must adhere to a 10-character naming
convention to successfully be read in (2-letter capitalized state
abbreviation followed by YYYYMMDD); if files were submitted
with a 5-digit format or lowercase state abbreviation, run fileRename() first.
In the example below, we will use the default read_hip()
settings to read in all of the fake HIP data files from an internal
package directory (read more in the example
data section). All other examples in this vignette will use
C:/ drive examples for clarity.
## Time to read in 49 files: 0 sec
The read_hip() function allows data to be read in
for:
- All states (e.g.,
state = NA, the default). - A specific state (e.g.,
state = "DE"). - A specific download (e.g.,
season = FALSE, the default;pathmust be to the download’s subdirectory). - An entire season (e.g.,
season = TRUE, andpathis to the directory containing all download subdirectories).
Use unique = TRUE to read in a frame without exact
duplicates, or unique = FALSE to read in all registrations
including exact duplicates. Important: the
record_key field is not created for
unique = FALSE, which is required in some following steps
(e.g., duplicateFix(), proof()).
Data are read in an expected format beginning with the
title field and ending with email field.
Additional fields are created by read_hip():
dl_state, dl_date, source_file,
dl_cycle, dl_key, record_key.
The read_hip() function does NOT read in data:
- From subfolders named
"hold","permit", or"lifetime". - Permit and lifetime HIP registrations must be read in and processed in a different workflow than the one outlined in this vignette.
qualityMessages
The qualityMessages() function was designed to find
issues in HIP registrations before data get too far in the
pre-processing pipeline. It requires this season’s start year to be
provided to the year parameter (e.g.,
year = 2026 for the 2026-2027 hunting season). Messages are
printed the console for the user to determine the level of severity and
what action to take next.
qualityMessages(raw_data, 2026)## Bad first name values detected.
## # A tibble: 10 × 4
## source_file file_size n_bad prop_bad
## <chr> <int> <int> <dbl>
## 1 /FL20260904.txt 225 6 0.027
## 2 /MS20260808.txt 226 6 0.027
## 3 /MI20260831.txt 225 5 0.022
## 4 /NY20260905.txt 227 5 0.022
## 5 /SD20260903.txt 226 5 0.022
## 6 /VA20260827.txt 227 5 0.022
## 7 /AR20260907.txt 226 4 0.018
## 8 /DE20260813.txt 225 4 0.018
## 9 /OK20260823.txt 226 4 0.018
## 10 /OR20260807.txt 225 4 0.018
## Bad last name values detected.
## # A tibble: 16 × 4
## source_file file_size n_bad prop_bad
## <chr> <int> <int> <dbl>
## 1 /KY20260821.txt 226 8 0.035
## 2 /CT20260901.txt 225 7 0.031
## 3 /CO20260827.txt 225 6 0.027
## 4 /AL20260803.txt 225 5 0.022
## 5 /MS20260808.txt 226 5 0.022
## 6 /WY20260822.txt 225 5 0.022
## 7 /AZ20260816.txt 226 4 0.018
## 8 /IA20260906.txt 225 4 0.018
## 9 /KS20260901.txt 225 4 0.018
## 10 /MO20260813.txt 226 4 0.018
## 11 /NC20260829.txt 225 4 0.018
## 12 /ND20260828.txt 228 4 0.018
## 13 /NE20260815.txt 225 4 0.018
## 14 /TN20260901.txt 225 4 0.018
## 15 /UT20260815.txt 225 4 0.018
## 16 /VA20260827.txt 227 4 0.018
## Bad suffix values detected.
## # A tibble: 49 × 4
## source_file file_size n_bad prop_bad
## <chr> <int> <int> <dbl>
## 1 /ME20260815.txt 225 128 0.569
## 2 /NC20260829.txt 225 122 0.542
## 3 /WA20260816.txt 226 122 0.54
## 4 /AL20260803.txt 225 121 0.538
## 5 /AZ20260816.txt 226 121 0.535
## 6 /NV20260810.txt 225 120 0.533
## 7 /WI20260824.txt 225 120 0.533
## 8 /LA20260804.txt 225 119 0.529
## 9 /MA20260804.txt 227 119 0.524
## 10 /NE20260815.txt 225 118 0.524
## # ℹ 39 more rows
## Bad state values detected.
## # A tibble: 49 × 4
## source_file file_size n_bad prop_bad
## <chr> <int> <int> <dbl>
## 1 /CT20260901.txt 225 20 0.089
## 2 /OH20260903.txt 227 20 0.088
## 3 /LA20260804.txt 225 19 0.084
## 4 /NM20260814.txt 227 19 0.084
## 5 /WV20260820.txt 226 18 0.08
## 6 /ME20260815.txt 225 17 0.076
## 7 /NV20260810.txt 225 17 0.076
## 8 /AR20260907.txt 226 16 0.071
## 9 /KS20260901.txt 225 16 0.071
## 10 /KY20260821.txt 226 16 0.071
## # ℹ 39 more rows
## Bad zip code values detected.
## # A tibble: 49 × 4
## source_file file_size n_bad prop_bad
## <chr> <int> <int> <dbl>
## 1 /ME20260815.txt 225 145 0.644
## 2 /MA20260804.txt 227 145 0.639
## 3 /AK20260813.txt 225 143 0.636
## 4 /NE20260815.txt 225 140 0.622
## 5 /CT20260901.txt 225 138 0.613
## 6 /WV20260820.txt 226 138 0.611
## 7 /CO20260827.txt 225 137 0.609
## 8 /FL20260904.txt 225 137 0.609
## 9 /OR20260807.txt 225 137 0.609
## 10 /NM20260814.txt 227 138 0.608
## # ℹ 39 more rows
## Error: 112 records detected with a value other than 2 for hunt_mig_birds.
## # A tibble: 45 × 3
## source_file hunt_mig_birds n
## <chr> <chr> <int>
## 1 /AK20260813.txt 1 2
## 2 /AL20260803.txt 1 1
## 3 /AZ20260816.txt 1 1
## 4 /CA20260808.txt 1 4
## 5 /CO20260827.txt 1 1
## 6 /CT20260901.txt 1 3
## 7 /DE20260813.txt 1 3
## 8 /FL20260904.txt 1 4
## 9 /GA20260831.txt 1 4
## 10 /IA20260906.txt 1 3
## # ℹ 35 more rows
## Error: 11 records have a '0' in every bag field; these records will be filtered out.
## # A tibble: 11 × 2
## source_file record_key
## <chr> <chr>
## 1 /CT20260901.txt record_1468
## 2 /GA20260831.txt record_2182
## 3 /ND20260828.txt record_5884
## 4 /ND20260828.txt record_5900
## 5 /OR20260807.txt record_7966
## 6 /OR20260807.txt record_8121
## 7 /OR20260807.txt record_8124
## 8 /WA20260816.txt record_10185
## 9 /WA20260816.txt record_10192
## 10 /WA20260816.txt record_10300
## 11 /WV20260820.txt record_10720
## Error: 1 test records detected; these records will be filtered out.
## # A tibble: 1 × 4
## source_file record_key firstname lastname
## <chr> <chr> <chr> <chr>
## 1 /WA20260816.txt record_10181 TEST TEST
## Error: 1 in-line permit records from OR and/or WA do not contain 2 for hunt_mig_birds; they will be edited.
## # A tibble: 1 × 6
## source_file hunt_mig_birds band_tailed_pigeon brant seaducks n
## <chr> <chr> <chr> <chr> <chr> <int>
## 1 /WA20260816.txt 1 0 0 2 1
## Error: 1702 records with non-zero bag values for permit species from permit file states; they will be edited.
## # A tibble: 18 × 4
## source_file spp strata n
## <chr> <chr> <chr> <int>
## 1 /CO20260827.txt cranes 1 97
## 2 /CO20260827.txt cranes 2 128
## 3 /KS20260901.txt cranes 2 112
## 4 /MN20260904.txt cranes 2 114
## 5 /MT20260817.txt cranes 2 113
## 6 /ND20260828.txt cranes 2 102
## 7 /ND20260828.txt cranes 9 1
## 8 /NM20260814.txt cranes 1 110
## 9 /NM20260814.txt cranes 2 117
## 10 /OK20260823.txt cranes 2 110
## 11 /OK20260823.txt cranes 9 1
## 12 /TX20260903.txt cranes 2 111
## 13 /WY20260822.txt cranes 1 69
## 14 /WY20260822.txt cranes 2 76
## 15 /CO20260827.txt band_tailed_pigeon 2 114
## 16 /NM20260814.txt band_tailed_pigeon 1 114
## 17 /NM20260814.txt band_tailed_pigeon 2 113
## 18 /UT20260815.txt band_tailed_pigeon 2 100
## Error: 3 registrations are missing critical combinations of PII (making up >10% of a file and/or >100 records).
## # A tibble: 3 × 3
## dl_state n proportion
## <chr> <int> <dbl>
## 1 CT 26 0.12
## 2 NM 22 0.1
## 3 OH 22 0.1
## Error: High non-resident proportions.
## # A tibble: 49 × 4
## source_file nonresident_n nonresident_prop contributing
## <chr> <int> <chr> <chr>
## 1 /ID20260904.txt 217 95.6% CA (6.2%), TX (5.7%), PA (5.3…
## 2 /RI20260905.txt 214 95.1% NY (6.7%), PA (6.7%), TX (6.2…
## 3 /IN20260826.txt 214 94.7% NY (5.8%), CA (5.3%), TX (4.4…
## 4 /VT20260831.txt 214 94.7% CA (5.8%), NY (5.3%), VA (4.9…
## 5 /NE20260815.txt 213 94.7% NY (6.7%), TX (5.3%), CA (5.3…
## 6 /ND20260828.txt 215 94.3% NY (5.3%), CA (4.8%), MA (3.9…
## 7 /SD20260903.txt 213 94.2% CA (8%), TX (5.3%), NY (4.9%)
## 8 /OK20260823.txt 213 94.2% CA (6.6%), KY (4.9%), NY (4%)
## 9 /MD20260829.txt 212 94.2% IA (6.2%), NY (4.9%), WV (4%)
## 10 /OR20260807.txt 212 94.2% CA (6.2%), VA (5.8%), NJ (4%)
## # ℹ 39 more rows
The qualityMessages() function returns a message
if:
-
NAvalues are detected in one or more ID fields (firstname,lastname,state,birth_date) for >10% of a file and/or >100 registrations. - All emails are missing from a file.
- Test records are found.
- Any registration has a
0in every bag field. - Any registration has an
NAin every bag field. - Any registration contains a bag value that is not a 1-digit number.
- For presumed solo permit did-not-hunts; if any Oregon or Washington
hunt_mig_birdsregistration value does not equal2when non-permit bags are0and one ofband_tailed_pigeon,brant, and/orseaduckis2. - Any
registration_yris not equal toREF_CURRENT_SEASONorREF_CURRENT_SEASON + 1. - Any file contains a
statevalue not equal todl_statethat makes up 10% or more of the file’s registrations. - Any inter-state duplicates are detected.
glyphCheck
During pre-processing, R may throw an error that says something like
“invalid UTF-8 byte sequence detected”. The error usually includes a
field name but no other helpful information. The
glyphCheck() function identifies values containing
non-UTF-8 glyphs/characters and prints them with the source file in the
console so they can be edited.
glyphCheck(raw_data)## All characters are UTF-8
shiftCheck
Find and print any rows that have a line shift error with
shiftCheck().
shiftCheck(raw_data)## No line shifts detected.
clean
The clean() function performs data cleaning and filters
out bad registrations.
clean_data <- clean(raw_data)## A total of 919 2s converted to 0s for permit file states:
## # A tibble: 9 × 3
## dl_state n spp
## <chr> <int> <chr>
## 1 CO 118 cranes
## 2 KS 100 cranes
## 3 MN 108 cranes
## 4 MT 105 cranes
## 5 ND 96 cranes
## 6 NM 110 cranes
## 7 OK 105 cranes
## 8 TX 107 cranes
## 9 WY 70 cranes
## A total of 305 2s converted to 0s for permit file states:
## # A tibble: 3 × 3
## dl_state n spp
## <chr> <int> <chr>
## 1 CO 107 band_tailed_pigeon
## 2 NM 106 band_tailed_pigeon
## 3 UT 92 band_tailed_pigeon
Registrations are dropped if:
- Any bag value is not a 1-digit number.
- Every bag field is
NAor0. - A test record is detected:
-
firstnameandlastnameare"TEST". -
lastnameis"INAUDIBLE". -
firstnameis one of:"INAUDIBLE","BLANK","USER","TEST", or"RESIDENT".
-
- There is missing contact/identification information:
-
NAvalue infirstname,lastname,state, orbirth_date. -
NAforaddressandemail. -
NAforemailandcityand/orzip.
-
Changes include:
-
firstname- Change to uppercase.
- If a suffix value is detected (e.g.,
JR,SR,1STto20TH, and1to20in Roman numerals, excludingXVIII) delete it. - Delete white space around string.
-
lastname- Change to uppercase.
- If a suffix value is detected (e.g.,
JR,SR,1STto20TH, and1to20in Roman numerals, excludingXVIII) delete it. - Delete white space around string.
-
suffix- Change to uppercase.
- If a suffix is detected in
firstnameorlastname, replace thesuffixwith that value. Values that are searched for includeJR,SR,1STto20TH, and1to20in Roman numerals, excludingXVIII. - Periods and commas are deleted.
-
zip- Remove ending hyphen from zip codes with only 5 digits.
- Remove ending
0from zip codes with 10 digits. - Insert a hyphen in continuous 9-digit zip codes.
- Insert a hyphen in 9-digit zip codes with a middle space.
- Delete trailing
-0000and-____.
-
hunt_mig_birds- For Oregon and Washington, if a registration’s
hunt_mig_birdsvalue is0and ifband_tailed_pigeon,brant, orseaduckis2, changehunt_mig_birdsvalue from0to2.
- For Oregon and Washington, if a registration’s
-
band_tailed_pigeon- If any permit file states submitted a
2forband_tailed_pigeon, change the2to a0.
- If any permit file states submitted a
-
cranes- If any permit file states submitted a
2forcranes, change the2to a0.
- If any permit file states submitted a
In addition to the changes listed above, the internal
zipCheck() function returns a message from
clean() if zip codes are detected that do not correspond to
provided state of residence for >10% of a file and/or >100
registrations.
issueCheck
The issueCheck() function assesses the validity of a
registration’s issue_date value. It does this based on the
current hunting season’s HIP issue window start date and end date (not
season open and close dates) for the state in which the registration was
issued. A plot is automatically returned for past and future
registrations. The plot is skipped by default, so to make one, specify
plot = TRUE.
current_data <- issueCheck(clean_data, year = 2026, plot = FALSE)Criteria evaluated and changes made:
-
Past registration: a registration’s
issue_dateis before the start date of HIP issuance for its state; these registrations are filtered out. -
Overlap registration: registrations from 2-season
states with an
issue_datebetween the start of next season’s issue start and this season’s issue end fall into the"overlap"category; these registrations are assigned"current"ifregistration_yris equal toyearor"future"ifregistration_yrisyear + 1. -
Current registration: a registration’s
issue_datefalls between the start and end dates of HIP issuance for its state;registration_yris overwritten withyear. -
Future registration: a registration has an
issue_dateafter the last day of issuance for its state for this season AND theissue_datefalls between the projected issue start and end dates for next season; theregistration_yris changed toyear + 1. -
Invalid registration: a registration’s
issue_datedoes not fall in the issue window for this season or next season; the date falls in the gap between issue windows. These registrations are filtered out. -
Bad
issue_date: a registration’sissue_datecannot be evaluated, likely because it’s formatted incorrectly or is illogical; these registrations are filtered out. - Return message if any registration’s
issue_datefalls after the file was submitted.
duplicateFinder
The duplicateFinder() function finds hunters that have
more than one registration. Registrations are grouped by
firstname, lastname, state,
birth_date, dl_state, and
registration_yr to identify unique hunters. If the same
hunter has 2 or more registrations, the fields that are not identical
are counted and summarized.
Plot the duplicates with duplicatePlot().
duplicateFinder(current_data)## There are 38 registrations with duplicates; 76 total duplicated records.
## # A tibble: 2 × 2
## duplicate_field n
## <chr> <int>
## 1 bag 14
## 2 issue_date 24
duplicateFix
We sometimes receive multiple HIP registrations per person which must
be resolved by duplicateFix(). Duplicates are identified
when more than one registration has the same firstname,
lastname, state, birth_date,
dl_state, and registration_yr. Only 1 HIP
registration per hunter per state can be kept. For in-line permit states
(WA, OR), permit records are submitted separately from HIP
registrations. Multiple permits are allowed. We differentiate
HIP and PMT records in in-line permit states
by checking the values in non-permit species fields
(ducks_bag, geese_bag, dove_bag,
woodcock_bag, coots_snipe, and
rails_gallinules); HIP registrations contain non-zero
values in those fields, but permit records always have 0
values.
deduplicated_data <- duplicateFix(current_data)To decide which HIP registration to keep from a group, we follow a series of logical steps.
For sea duck and brant states:
These states include: AK, CA, CT, DE, MA, MD, NC, NH, NJ, NY, RI, VA
- Keep registration(s) with the most recent issue date.
- Exclude registrations with all
1values or all0values in bag fields from consideration. - Keep any registrations that have a
2for eitherbrantorseaducks(for sea duck and brant states), or2forseaducks(Maine only). - If more than one registration remains, choose to keep one randomly.
For HIP registrations from in-line permit states and all other states:
These states include: AL, AZ, AR, CO, FL, GA, ID, IL, IN, IA, KS, KY, LA, MI, MN, MS, MO, MT, NE, NV, NM, ND, OH, OK, OR, PA, SC, SD, TN, TX, UT, VT, WA, WV, WI, WY
- Keep registration(s) with the most recent issue date.
- Exclude registrations with all
1values or all0values in bag fields from consideration, if possible. - If more than one registration remains, choose to keep one randomly.
A new field called record_type is added to the data
after the above deduplication process. Every HIP registration is labeled
HIP. Registrations from in-line permit states (Washington
and Oregon) can be labeled HIP or PMT.
bagCheck
The bagCheck() function searches for values that are not
typical or expected in all species group fields. If a value outside of
the normal range is detected, an output tibble is created. Each row in
the output contains the state, species, unusual value, and a list of the
normal values we would expect. If a value for a species group is
provided that doesn’t match anything in our records, the output will
show NA values in the expected_bag_value
column. These species do not have a hunting season in the reported
states.
In-line permit records are not included in this check, to prevent
ducks_bag = "0" values from popping up when we know those
don’t count.
bagCheck(deduplicated_data)## # A tibble: 88 × 6
## dl_state spp bad_bag_value expected_bag_value n proportion
## <chr> <chr> <chr> <chr> <int> <chr>
## 1 MD woodcock_bag 9 2, 3, 5 1 0%
## 2 FL woodcock_bag 9 1, 2, 3, 5 1 0%
## 3 GA dove_bag 9 1, 2, 3, 5 1 0%
## 4 KS dove_bag 9 1, 2, 3, 5 1 0%
## 5 KY dove_bag 9 1, 2, 3, 5 1 0%
## 6 KY woodcock_bag 9 1, 2, 3, 5 1 0%
## 7 NJ woodcock_bag 9 1, 2, 3, 5 1 0%
## 8 MT coots_snipe 9 1, 2, 3, 4, 5 1 0%
## 9 NC ducks_bag 9 1, 2, 3, 4, 5 1 0%
## 10 NC dove_bag 9 1, 2, 3, 4, 5 1 0%
## # ℹ 78 more rows
proof
We run proof() to check the data for errors. The
season’s start year must be supplied to the year parameter
(e.g., year = 2026 for the 2026-2027 hunting season). This
helps to check values in registration_yr and
birth_date. The output of the proof() function
contains a new field called errors.
The fields proofed for errors include title,
firstname, middle, lastname,
suffix, address, city,
state, zip, birth_date,
hunt_mig_birds, registration_yr, and
email. When values are found to be irregular, the field
name(s) is/are added to errors. The errors
field pastes all of the registration’s errors together as a
hyphen-delimited string; a row can have zero errors (e.g.,
NA), one error (e.g., "zip"), two errors
(e.g., "suffix-zip"), three errors (e.g.,
"suffix-address-zip"), etc; up to all 12 field names.
Note that no actual corrections or data changes take place as a
result of the proof() function.
proofed_data <- proof(deduplicated_data, year = 2026)What gets flagged in the errors field and why:
-
"title"- If
titleis not one of:NA,0,1or2. - If common first names are assigned the wrong title.
- If
-
"firstname"-
firstnamecontains anything other than letters, apostrophe(s), space(s), and/or hyphen(s). -
firstnamecontains less than 2 letters. -
firstnamecontains"AKA".
-
-
"middle"ifmiddleis not exactly 1 letter orNA. -
"lastname"-
lastnamecontains anything except letters, apostrophe(s), space(s), hyphen(s), and/or period(s). -
lastnamecontains less than 2 letters.
-
-
"suffix"-
suffixshould be one of:I, II, III, IV, V, VI, VII, VIII, IX, X, XI, XII, XIII, XIV, XV, XVI, XVII, XIX, XX, 1ST, 2ND, 3RD, 4TH, 5TH, 6TH, 7TH, 8TH, 9TH, 10TH, 11TH, 12TH, 13TH, 14TH, 15TH, 16TH, 17TH, 18TH, 19TH, 20TH, JR, SR. - Note that
XVIIIis excluded (exceeds 4 character limit).
-
-
"address"ifaddresscontains a|, tab, or non-UTF8 character. -
"city"-
citycontains anything other than letters, space(s), hyphen(s), apostrophe(s), and/or period(s). -
citycontains less than 3 letters.
-
-
"state"- If
stateis not contained in the following list of abbreviations for US states and territories and Canadian provinces and territories: AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, WV, WI, WY, DC, AS, GU, MP, PR, VI, UM, MH, FM, PW, AA, AE, AP, AB, BC, MB, NB, NL, NS, NT, NU, ON, PE, PQ, QC, SK, YT.
- If
-
"zip"- If a registration’s
addressdoesn’t have azipthat should be in their reported addressstateof residence. - If the fist 5 digits of
zipare not inREF_ZIP_CODE$zipcodeat all.
- If a registration’s
-
"birth_date"- If the date is formatted poorly (per
REGEX_DATE_FORMAT). - If the date can’t be parsed by
lubridate::mdy(). - If the date is later than today.
- If the date is earlier than
09/01/(REF_CURRENT_SEASON - 100).
- If the date is formatted poorly (per
-
"hunt_mig_birds"if not equal to1or2. -
"registration_yr"if not equal to the current season. -
"email"-
emaildoes not match universally accepted email regex (seeREGEX_EMAIL). -
emailis obfuscative:- e.g.,
none@gmail.com,nottoday@gmail.com,fake.fake@gmail.com,n@a.com,x@y.com,brian@na.com, etc (seeREGEX_EMAIL_OBFUSCATIVE_LOCALPART,REGEX_EMAIL_OBFUSCATIVE_DOMAIN, andREGEX_EMAIL_OBFUSCATIVE_ADDRESS). - A repeated character is detected, e.g.
aaa@a.com(seeREGEX_EMAIL_REPEATED_CHAR). - The domain is
tpwd.texas.govor some variation (seeREGEX_EMAIL_OBFUSCATIVE_TPWD). - A
walmart.comdomain is preceded by only numbers (seeREGEX_EMAIL_OBFUSCATIVE_WALMART).
- e.g.,
-
emailis longer than 100 characters. - A common domain name (e.g., gmail, yahoo) has a common typo.
- A common domain name doesn’t have a matching top-level domain (e.g.,
gmail.netorhotmail.gov). - The address has a bad top-level domain (e.g.,
.comcom,.ccom, etc). - The email is missing a top-level domain.
- The top-level domain period is missing (e.g.,
gmailcom).
-
correct
Data can be corrected by running the correct() function.
Users must provide the season’s start year to the year
parameter (e.g., year = 2026 for the 2026-2027 hunting
season).
corrected_data <- correct(proofed_data, year = 2026)Changes and corrections include:
-
titleis changed toNAif"title"is in theerrorsfield. -
middleis changed toNAif"middle"is in theerrorsfield. -
suffixis changed toNAif"suffix"is in theerrorsfield. -
email- Add endings to common domains if missing (e.g.
@gmailwould become@gmail.com), for domains including:-
gmail, yahoo, hotmail, aol, icloud, comcast, outlook, sbcglobal, att, msn, live, bellsouth, charter, ymail, me, verizon, cox, earthlink, protonmail, pm, mail, duck, ducks.
-
- Add top-level domain period(s) if:
- Missing before
com, net, edu, gov, org. - Missing in
navymil, mailmil, armymil; add multiple periods if missing tousnavymil, usafmil, usarmymil, usacearmymil.
- Missing before
- Add endings to common domains if missing (e.g.
-
errorsis updated by re-runningproof()inside ofcorrect().
Part B: Data Visualization and Tabulation
duplicatePlot
Plot duplicates with duplicatePlot().
duplicatePlot(current_data)## There are 38 registrations with duplicates; 76 total duplicated records.

Figure 1. Plot of types of duplicates.
errorPlotFields
The errorPlotFields() function can be run on all
states…
errorPlotFields(proofed_data, loc = "all", year = 2026)
Figure 2. Plot of all location’s errors by field name.
… or it can be limited to just one.
errorPlotFields(proofed_data, loc = "LA", year = 2026)
Figure 3. Plot of Louisiana’s errors by field name.
It is possible to add any ggplot2 components to these
plots. A plot can be altered with facet_wrap using either
dl_cycle or dl_date. The example below
demonstrates how this package’s functions can interact with
tidyverse and shows an example of an
errorPlotFields with facet_wrap (using a
subset of 4 download cycles).
errorPlotFields(
hipdata2025 |>
filter(str_detect(dl_cycle, "0800|0901|0902|1001")),
year = 2026) +
theme(
axis.text.x = element_text(angle = 90, vjust = 0, hjust = 1),
legend.position = "bottom") +
facet_wrap(~dl_cycle, ncol = 2)errorPlotStates
The errorPlotStates() function plots error proportions
per state. A threshold value must be set to only view states above a
certain proportion of error (in the example below,
threshold = 0.05 indicates an error tolerance of 5%). Bar
labels are error counts.
errorPlotStates(proofed_data, threshold = 0.05)
Figure 4. Plot of proportion of error by state.
errorPlotDL
This function should not be used unless you want to visualize an
entire season of data. The errorPlotDL() function plots
proportion of error per download cycle over the course of the hunting
season. Location may be specified with the loc parameter to
see a particular state over time.
errorPlotDL(hipdata2025, loc = "MI")
errorPlot_dl example output
pullErrors
The pullErrors() function can be used to view all of the
actual values that were flagged as errors in a particular field. In this
example, we find that the suffix field contains several
values that are not accepted.
pullErrors(proofed_data, field = "suffix")## [1] "DDS" "MD" "PHD" "DVM"
Running pullErrors() on a field that has no errors will
return a message.
pullErrors(proofed_data, field = "dove_bag")## Success! All values are correct.
errorTable
The errorTable() function returns error data as a
tibble, which can be assessed as needed, or exported to create records
of download cycle errors. The basic function reports errors by both
location and field.
errorTable(proofed_data)## # A tibble: 254 × 3
## dl_state error error_count
## <chr> <chr> <int>
## 1 AK email 3
## 2 AK firstname 1
## 3 AK lastname 2
## 4 AK suffix 71
## 5 AK zip 131
## 6 AL email 1
## 7 AL firstname 1
## 8 AL lastname 5
## 9 AL suffix 79
## 10 AL zip 122
## # ℹ 244 more rows
Errors can be reported by only location by turning off the
field parameter.
errorTable(proofed_data, field = "none")## # A tibble: 49 × 2
## dl_state error_count
## <chr> <int>
## 1 AK 208
## 2 AL 208
## 3 AR 200
## 4 AZ 212
## 5 CA 203
## 6 CO 214
## 7 CT 184
## 8 DE 200
## 9 FL 213
## 10 GA 194
## # ℹ 39 more rows
Errors can be reported by only field by turning off the
loc parameter.
errorTable(proofed_data, loc = "none")## # A tibble: 6 × 2
## error error_count
## <chr> <int>
## 1 birth_date 25
## 2 email 154
## 3 firstname 105
## 4 lastname 130
## 5 suffix 3363
## 6 zip 5963
Location can be specified using one of the 49 contiguous state abbreviations.
errorTable(proofed_data, loc = "CA")## # A tibble: 5 × 3
## dl_state error error_count
## <chr> <chr> <int>
## 1 CA email 2
## 2 CA firstname 1
## 3 CA lastname 2
## 4 CA suffix 69
## 5 CA zip 129
Field can be specified (one of: "all",
"none", "title", "firstname",
"middle", "lastname", "suffix",
"address", "city", "state",
"zip", "birth_date",
"hunt_mig_birds", "registration_yr",
"email").
errorTable(proofed_data, field = "suffix")## # A tibble: 1 × 2
## error error_count
## <chr> <int>
## 1 suffix 3363
Total errors for a location can be pulled.
errorTable(proofed_data, loc = "CA", field = "none")## # A tibble: 1 × 2
## dl_state total_errors
## <chr> <int>
## 1 CA 203
Total errors for a field in a particular location can be pulled.
errorTable(proofed_data, loc = "CA", field = "dove_bag")## No errors in dove_bag for CA.
Part C: Writing Data and Reports
write_hip
After the data have been corrected, the data are ready to be written
out. Use write_hip() to do final processing to the data,
which includes 1) adding in FWS strata and 2) setting NA
values to blank strings. The write_hip() function will fail
if there is an NA in dl_state or if there is
an NA in dl_date. It will also fail if any
value exceeds the database character limits (see
failWidths() helper function).
If split = FALSE, the final table will be saved as a
single .csv to your specified path. If
split = TRUE (default), one .csv file per each
input .txt source file will be written to the specified
directory.
The type argument helps apply additional checks for
permits; supply "HIP" for HIP data from a regular download,
or "CR" or "BT" for separate permit file
records. If type = "HIP", writing data will fail if there
are more record_type = "PMT" records in the
corrected_data than record_type = "HIP"
records.
write_hip(
corrected_data,
path = "C:/HIP/processed_data/",
type = "HIP",
split = TRUE)writeReport
The writeReport() function can be used to automatically
generate an R markdown document with figures, tables, and summary
statistics. This can be done at the end of a download cycle.
writeReport(
raw_path = "C:/HIP/DL0901/",
temp_path = "C:/HIP/corrected_data",
year = 2026,
dl = "0901",
dir = "C:/HIP/dl_reports",
file = "DL0901_report")